Generate code from the spec, not the spec from the code
Here is the failure pattern every team using a coding agent has hit by week two. The agent builds the screen fast, wires the form, posts to the create endpoint — and the integration is wrong. limit became pageSize. The list came back as data.records in the prompt's imagination and data.items in reality. Money was generated as a float. The agent did not make a stupid mistake; it made the most likely guess from a contract it had never seen.
The fix is not a sharper prompt. It is to make the specification the thing the agent reads — and then to verify generated code against that specification with executable scenarios. Code generation works when it flows from the spec, not when the spec has to be reconstructed from generated code afterward.
Give the agent the contract, not a summary of it
There are two ways to hand a coding agent the OpenAPI document, and both are already supported by the same workspace that produced it. Point the agent at the spec file directly — it is a local YAML document in version control, like any other source of truth — or serve the workspace over MCP and let the agent discover operations, schemas, and examples as tools. The MCP route wins on larger surfaces: instead of pasting three thousand lines into context, the agent queries the operation it is implementing, gets the exact request schema, response schema, status codes, and auth requirements, and cannot paraphrase a field name in the process.
The spec also travels with the decisions that never make it into a generated type: the integer-cents convention, the idempotency requirement, the shared error envelope. Those descriptions exist because the design session forced them into the open before any code existed. Regenerating types preserves field names; serving the contract preserves intent.
Generate in the order that fails cheaply
Asking an agent to "implement the reservations API" produces forty endpoints of plausible code in one pass, which is forty endpoints of unverified assumptions. Slice vertically, and generate in this order:
| Order | Artifact | Why first |
|---|---|---|
| 1 | Types and DTOs from the schemas | Mechanical, high-volume, and the foundation everything else compiles against. |
| 2 | Boundary validation | Wrong input should fail at the framework edge with the documented error envelope, never reach the service layer. |
| 3 | One happy path end to end | Create hold → 201, wired through the real database and the real clock. |
| 4 | The documented failure cases | 409 on a duplicate, 422 on a bad body, idempotent replay returning the original resource. |
| 5 | The remaining operations | Copied from a proven slice instead of five independent guesses. |
The vertical slice matters more than it looks. A single operation working against the real stack settles the routing conventions, dependency injection, transaction boundaries, and response serialization for every route that follows. It also gives the verification loop something real to hit.
Verify with the scenarios, then feed failures back as evidence
The same scenario suite written during design — create, idempotent replay, duplicate conflict, event confirmation — now runs against localhost instead of the mock. This is the moment generation becomes trustworthy. When the agent's code returns a bare string on 409 instead of the error envelope, do not describe the bug to it in prose. Hand it the failed scenario: the request, the expected schema, the actual response. That is evidence, not opinion, and agents fix evidence far more reliably than they fix vibes.
Run the suite in CI from the run-host guide so a pull request that drifts from the contract fails the build. The common failure modes — renamed fields, missing status codes, responses that only work on the happy path — are exactly what the integration-is-wrong post catalogs, and every one of them is a scenario assertion away.
Close the loop with a rescan
Code and spec drift even under disciplined teams; under AI-assisted velocity they drift faster. The scanner that reconstructs a spec from an existing codebase runs in the other direction on purpose. After a sprint of generated code, rescan the project, diff the discovered routes against the design spec, and reconcile the difference:
- A route exists in code but not the spec — either the agent invented an endpoint, or the spec missed a requirement. One of those is a bug.
- A route exists in the spec but not in code — unfinished work, now visible before release night.
- A parameter or response field differs — the contract broke somewhere, and the sidecar diff shows exactly where.
Design-first and code-scanning are two directions of the same loop, and using both is what keeps either one honest. The spec is authored up front with AI drafting and humans reviewing; code is generated against it; scenarios verify behavior; rescans detect drift. Switching models or agents halfway through changes nothing, because none of the contract lives inside the assistant — it lives in the file, which is also why changing AI tools never means reintroducing your API.
What not to do
- Do not generate the whole surface in one prompt and call it done; untested generated code is just debt with a new author.
- Do not let the agent invent fields "for convenience." If the schema is wrong, change the spec, review the diff, and regenerate.
- Do not paste a flattened summary of the API into chat when the agent can query the document over MCP.
- Do not treat a passing build as a passing contract — run the scenarios against the running service.
The local-first workspace is doing the unglamorous work underneath: one OpenAPI file as source of truth, mock and scenarios derived from it, MCP serving it, and the scanner checking the code back against it. The AI does the drafting and the typing. The spec does the remembering. The scenarios decide what done means.