Watch an engineer use a coding agent against an unfamiliar API and you see the same ritual every session. They paste the docs URL, then the auth section, then an example response, then a correction: "pagination is cursor, not page", "money is integer cents", "the list wrapper is data.items". The agent writes plausible code, fails on the first run, and the engineer feeds it another screenshot. Context, assembled by hand, decays within the hour.

There is a structural fix, and it is not better prompts. It is giving the agent the contract as callable tools over MCP, generated from the same OpenAPI document the team maintains.

Why prose is the wrong transport

Dumping documentation into a chat has three failure modes that never go away:

  1. Staleness. The pasted text is a snapshot. The spec moved last sprint; the agent's context did not, and nothing warns anyone.
  2. Token cost paid repeatedly. Every conversation re-uploads the same 40 pages. Agents that re-read files each turn multiply it.
  3. No enforcement. Prose says "ids are UUIDs under data.id"; the agent writes resp.id anyway because nothing validates the call until it fails at runtime.

RAG pipelines over docs improve retrieval and do nothing for enforcement — the model still describes the call it is about to make, with all the usual hallucination surface.

What the MCP shape changes

An OpenAPI-to-MCP server turns each operation into a tool. The agent does not read that POST /v1/orders exists; it sees a create_order tool whose inputSchema requires plan_code and quantity, rejects a missing field before the HTTP call, and returns the parsed response. The differences are practical:

  • Discovery is structured. Tool names and JSON schemas are built for model selection; prose is built for humans skimming.
  • Validation happens pre-flight. A wrong type is a tool error the agent corrects immediately, not a 422 from staging the engineer has to read back to it.
  • Auth is centralized. Tokens live in the MCP server config, never in prompts or committed code.
  • The contract updates in one place. Ship a new spec version; every agent attached to the endpoint gets it with no re-pasting.
  • Streaming operations stay usable. An SSE operation documented with its event schema becomes a tool that opens the stream and returns the collected events instead of the agent pretending streams do not exist.

What the setup actually looks like

Local development uses a local server over the spec file:

{
  "mcpServers": {
    "billing-api": {
      "command": "npx",
      "args": ["@powerduck/openapi-to-mcp-server", "--spec", "/Users/me/work/billing.openapi.yaml"],
      "env": { "PD_API_TOKEN": "scoped-local-token" }
    }
  }
}

Point Cursor, Claude Code, or any MCP client at that config and the billing operations appear as tools. For shared environments — partner integrations, internal support tooling — the same spec is published as a hosted MCP endpoint with a per-user token, version pinned so an agent integration does not break on a spec release.

The token scoping deserves emphasis. Give the agent's token the minimum verbs: read-only for code-generation tasks, write scoped to a sandbox organization for anything that creates data. An agent that can browse the catalog should never hold a token that can delete accounts.

Where MCP genuinely wins

  • Repeated integration work against your own API. Every internal agent, codegen task, and support script gets the same typed surface.
  • Environments with many small operations. Agents navigate dozens of narrow tools better than one 400-page document.
  • Enforcement-heavy domains. Billing, identity, and compliance work where a malformed request is expensive — pre-flight schema validation pays for itself.
  • Fast-moving specs. When the contract changes weekly, anything copied into prompts is already wrong.

Where it does not

MCP is not a universal replacement for docs or SDKs, and selling it as one creates new problems:

  • One-off exploration. A developer evaluating the API wants narrative guides, concepts, and sequencing — tools assume you already know what you are doing. Keep the rendered docs.
  • Performance-critical or batch clients. An agent calling a tool per resource is fine; production code should use a generated SDK with connection pooling, retries, and typed models. MCP is for the agent's working loop, not the shipped runtime.
  • Operations that are not in the spec. MCP exposes exactly what the contract says. If the spec is a lie, the tools are lies — the OpenAPI document still has to be maintained, ideally as the source of design rather than a reverse-engineered afterthought.
  • Huge flat surfaces. 300 unprefixed tools degrade selection; expose tagged subsets per audience.

The honest division of labor

The mature setup in 2026 is three renderings of one OpenAPI document, each for a different consumer:

  1. Rendered documentation for humans learning the API.
  2. Generated SDKs for production code.
  3. An MCP server for agents doing integration and exploration work.

Maintain the spec once; derive all three. The failure mode to avoid is four teams independently describing the same endpoints — one in Markdown, one in the SDK, one in the MCP tool wrappers, one in the test fixtures. That is the exact duplication the contract was supposed to remove.

Powerduck serves the local MCP endpoint from the spec open in the workspace and publishes a hosted version with access control from Cloud; the demo shows the serve action on a sample document. The step-by-step conversion is in how to turn an OpenAPI spec into an MCP server.

What to read next: switching AI tools without reintroducing your API is about not locking the contract into one vendor's agent, and when AI writes your system, who defines done covers the verification side.