The gap between a demo agent and a reliable one is usually not model intelligence; it is output discipline. In a demo, the model calls get_weather({"city": "SF"}) and looks magical. In production it calls refund_payment({"payment_id": 7712, "amount": "49.00", "currency": "usd", "reason": "user asked nicely"}) — integer where you expect a string, string where you expect cents, an undeclared reason, and an extra field your handler silently ignores. Every one of those mismatches is a bug report that says "the AI is unreliable" when the real cause is an unconstrained output contract.

Structured outputs fix this at the model layer, and JSON Schema is the language every provider converged on. If you maintain an OpenAPI document, you already own most of the schemas involved.

What structured outputs actually guarantees

Providers implement this under different names — OpenAI's structured outputs and function calling strict mode, Anthropic's tool-use input schemas, Gemini's responseSchema — but the guarantee is the same shape: given a schema the provider can enforce, the response is guaranteed to validate against it. Invalid JSON, missing required fields, and types that do not match are eliminated at decoding time rather than surfacing in your application.

That guarantee is deliberately narrow. It does not promise the values are correct (the model can still refund the wrong payment), only that they are valid. Validation is exactly the layer you should never have had to hand-write in natural language; correctness still requires good tool design, confirmation gates, and tests.

The strict subset you must design within

Providers enforce structured outputs by constraining the JSON Schema dialect they accept. The rules are consistent across the major platforms and easy to internalize:

  • Every object must list "additionalProperties": false and enumerate all properties.
  • Every property must appear in required; optionality is expressed with a union including null, not by omission.
  • Types are explicit; use type: ["string", "null"] for nullable fields.
  • $ref is supported against $defs (or definitions), which keeps schemas DRY.
  • Avoid unsupported keywords (format is often advisory, not enforced; avoid patternProperties, conditional schemas, and tuple-heavy arrays unless your provider documents them).

A strict refund schema looks like this:

{
  "type": "object",
  "additionalProperties": false,
  "properties": {
    "payment_id": { "type": "string" },
    "amount_cents": { "type": ["integer", "null"], "minimum": 1 },
    "currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] },
    "reason": {
      "type": "string",
      "enum": ["customer_request", "duplicate", "fraud"]
    },
    "notify_customer": { "type": "boolean" }
  },
  "required": ["payment_id", "amount_cents", "currency", "reason", "notify_customer"]
}

There is nowhere for an invented field to land, no ambiguity about whether the amount is a string, and a closed enum for the reason instead of a prose judgment the support team later has to parse.

Enums and unions carry the semantics

Models are remarkably good at mapping messy human intent onto closed value sets, and remarkably bad at inventing consistent open-ended codes. Push decisions into schemas wherever a value set is finite: statuses, categories, sort orders, time windows. When a concept is genuinely open (a customer-facing note), keep it a string and constrain the rest of the shape.

For nullable optionality, prefer the union form over omitting the key. A consistent object shape simplifies both the model's job and your handler: there is no difference between "key absent" and "key present with null" to test.

Reuse the schemas you already publish

If the tool wraps an HTTP API, its arguments are the API's request body, path parameters, and query parameters. Those are already modeled in OpenAPI under components.schemas. Duplicating them into tool definitions creates two contracts that drift the first time someone edits one. Generate the tool input schema from the operation:

paths:
  /v1/payments/{id}/refunds:
    post:
      operationId: refundPayment
      parameters:
        - in: path
          name: id
          required: true
          schema: { type: string }
      requestBody:
        required: true
        content:
          application/json:
            schema: { $ref: '#/components/schemas/RefundRequest' }

A build step can combine the path parameters and the request schema into one strict tool input object, namespacing $defs to avoid collisions across services. The same component schema then drives:

  • Server-side validation (the API itself).
  • Generated typed clients in TypeScript and other languages.
  • MCP tool input schemas for agents.
  • Documentation examples.

This is the same single-source-of-truth argument behind generating a TypeScript client from OpenAPI, applied to the model interface.

Validate anyway, and make errors teachable

Provider guarantees remove malformed output; they do not remove business-rule violations (refund over the captured amount, payment already refunded). Keep a server-side validation pass using the same schema, and return field-level, machine-readable errors so the model can self-correct in one turn:

{
  "type": "https://errors.example.com/validation-failed",
  "status": 422,
  "errors": [
    { "field": "amount_cents", "code": "above_captured_amount", "max": 4900 }
  ],
  "retryable": true
}

A model receiving that response retries with a corrected value. A model receiving "refund failed" guesses. Error design for this audience is covered in designing APIs for AI agents and the error format itself in RFC 9457 Problem Details.

Test the schema boundary like any other contract

  • Schema meta-validation: every tool input schema validates against the strict subset; reject PRs that introduce unsupported keywords, which providers either reject at registration or silently downgrade.
  • Representative prompts: run a fixed set of natural-language requests through the model in CI against recorded responses, asserting they validate. This catches description problems (the model cannot tell which field is the amount) without testing the model itself.
  • Enum coverage: each enum value is reachable from a realistic phrasing; if users say "cancel" and the enum only knows refund, the description or the enum needs work.

The payoff across tools, MCP, and clients

When OpenAPI components are the canonical schemas, strict structured outputs propagate everywhere an agent touches the platform: direct function calling in an application, MCP tools generated from the spec, and typed SDKs for human-written code all enforce the same shapes. Agents stop being a class of integration that needs bespoke parsing and retry logic, and "the AI sent bad arguments" shrinks from a daily failure category to a schema-regression caught in CI.

A local-first, spec-driven workspace is built around exactly this loop: design the schemas, generate strict tool definitions, and exercise them against mocks before the API exists. You can see a spec turned into tool-ready structures in the online demo, and the agent-facing API design principles that pair with strict schemas are summarized in when AI writes your system, who defines done.