I gave the assistant a one-paragraph brief — a marketplace for equipment rentals — and asked for the API. It proposed forty operations overnight. The first pass had the usual AI problems: rental, booking and order used interchangeably, money as a float in one endpoint and integer cents in another, nesting that changed path to path.

The second pass was buildable. Same model, same brief. What changed was the machinery between the model and the file. Every gate below is a real part of the design surface in Powerduck.

Gate 1: broad requests are staged, not dumped

For "build an e-commerce API" the assistant does not emit forty operations at once. It asks one clarifying question about the key decisions, emits a plan of the endpoints it intends to create, then proposes focused patches of two to five operations per card. The resource-naming drift died at the planning card, before a single schema existed.

Gate 2: patches are granular JSON Patch, never whole-operation rewrites

This is the gate I underestimated. When you refine an existing operation, the assistant does not resend the whole object (which would silently wipe every field it failed to echo). It sends small operations deep inside the document. Path segments are array elements:

// append one query parameter, nothing else changes
{ "op": "add",
  "path": ["paths", "/products", "get", "parameters", "-"],
  "value": { "name": "limit", "in": "query", "schema": { "type": "integer" } } }

// tighten one field in one response schema
{ "op": "replace",
  "path": ["paths", "/products", "get", "responses", "200",
           "content", "application/json", "schema",
           "properties", "total"],
  "value": { "type": "integer" } }

Arrays append with the "-" token instead of being resent wholesale, and parameters are keyed by (in, name) so the same parameter is never duplicated. When I said "make the customer email required," that one property changed and everything else in the operation stayed byte-for-byte intact.

Gate 3: read before edit

Before proposing anything, the model can call spec.overview, spec.listOperations, spec.getOperation and spec.getSchema. The host expects it to read an operation before editing it whenever the full shape is not visible. An assistant that has fetched the real schema cannot honestly invent a field — and the current document, not the model’s memory of an earlier patch, is always the authoritative state.

Gate 4: deterministic validation and card types

The model proposes; the host validates. Not every reply is a patch: question cards make me choose between options, validation cards report pass/warning/error checks, action cards offer to run a scenario, and data-table cards hand back concrete sample rows. Write actions always wait for confirmation.

Gate 5: drift rules while iterating

In focused mode only the active operation — plus schemas it explicitly references — can change; there are no opportunistic edits to sibling endpoints. Naming a different METHOD /path is treated as a deliberate task switch. Applied and rejected proposals are tracked, so a patch already in the document is not proposed twice.

The scorecard

Of the forty operations I kept thirty-four as proposed, merged four pairs, and redesigned two — the money representations, predictably. I reviewed cards in about an hour of evening time. The point is not that the AI was perfect; it is that every mistake was caught at the gate where it cost seconds instead of sprints.

Try the loop on your own spec with the designing guide — and keep the model on a short leash with focused cards.