A 200 is not done: make the AI walk the business scenario
The create-order endpoint returned 200. The AI marked the task complete.
You clicked through to payment and the order detail view could not find the order that had just been created. Then cancel turned out to accept a status that should never be cancellable.
A single successful request produces a strong feeling of completion, but users never experience one request. They experience a sequence: create, query, change state, confirm the result. As AI takes over more implementation, acceptance has to move from "it looks written" to "the scenario actually ran."
Hand over the verification process with the requirement
For a stripped-down order feature, settle the create, query, and cancel operations in the specification, then give the agent a task like this:
Design a test scenario from the current API specification: create an order, extract the order ID from the response, query the order with that ID, cancel it, and query the final state. Add sensible assertions at each step. Show me the scenario before running it.
The important part is the relationship between steps. The query uses the freshly created ID, not a hard-coded sample. The final assertion checks the state after cancellation, not merely whether the cancel call returned 200. Powerduck's dev MCP server runs ordered scenarios, extracts values between steps, and evaluates assertions, so those relationships survive into actual execution.
Turn a chat instruction into a scenario that runs again tomorrow
A scenario worth keeping spells out:
- which operation each step calls and with what parameters;
- which value is extracted from a response and where it goes next;
- what must hold at each step, and whether failure stops the run.
The scenario is statically checked first — every operation it references has to exist in the specification. After a test base URL and credentials are configured, it runs against the real service, showing each response and assertion result. Saved as a project file, it does double duty: today it accepts the AI's change; after a refactor it checks that the same business path still works. Feedback starts to accumulate. A failure leaves the steps, inputs, responses, and failed condition instead of a message that says "still broken."
Scenarios are deliberately kept out of the OpenAPI document itself — there are no embedded test keys in the contract — and the assistant sequences the work so endpoint patches are proposed and applied before the run card appears.
The agent gets something it can actually fix with
Say the create call succeeds and the detail query returns 404. The agent now has a trail to follow: did the create response include the right ID, is the query hitting the same environment, did the implementation actually persist the write? After the fix, the same scenario runs again with identical assertions, so the results are comparable.
That closes a loop that fits AI coding well:
read the spec → implement or change → run the scenario → inspect the failure evidence → fix again.
People keep the business judgments. Whether cancellation releases inventory is a rule the business has to state; three endpoint names cannot derive it.
"Done" should carry evidence
Scenario testing is not the whole test strategy. Concurrency, transactional consistency, and complex permissions still need their own methods, and anything that writes should target a proper test environment with a plan for cleaning up the data.
For day-to-day AI-assisted development, though, reliably walking the key business paths is far closer to the user's real experience than confirming that one endpoint answers. When the AI says it is finished, the useful follow-up is already obvious:
Which scenario did you run, what did each step verify, and what failed or stayed uncovered? List all of it.
The deliverable from AI coding can include both the code and the evidence that it works.