Scenario testing for REST APIs: writing user-journey tests from OpenAPI
Single-request testing answers a narrow question: does this endpoint, given these inputs, return this status right now? It cannot answer the question that actually breaks releases: can a user complete the job? A user never calls one endpoint. They create a resource, wait for it to become ready, attach something to it, act on it, and verify the result. Each request depends on the previous one's response. That chain is a scenario, and it is the level where integration bugs actually live.
Why request-level tests miss everything
A typical collection has forty green requests. Then production shows:
- The create endpoint returns an
idnested underdata.id, but the next request was written assumingidat the root. - A list endpoint paginates with a cursor the test never saw because fixtures returned three items.
- A status field transitions
pending → processing → ready, and the test assertedreadyimmediately. - An auth token minted for user A is used to read user B's resource and the server allowed it.
Every request in isolation returned 200. The journey was never tested. Scenario testing makes these dependencies explicit: variables extracted from one response become inputs to the next, and assertions run on the accumulated state.
The anatomy of a scenario
A scenario is a sequence of steps. Each step has a request, extraction rules, and assertions:
name: Create and cancel an order
steps:
- name: create order
request:
method: POST
path: /v1/orders
body:
plan_code: TEAM
quantity: 1
assert:
status: 201
json:
data.status: pending
extract:
order_id: data.id
- name: fetch the order
request:
method: GET
path: /v1/orders/{{order_id}}
assert:
status: 200
json:
data.id: "{{order_id}}"
- name: cancel
request:
method: POST
path: /v1/orders/{{order_id}}/cancel
assert:
status: 200
json:
data.status: canceledThe shape is deliberately boring: HTTP verbs, paths, JSONPath-ish assertions, and a template variable. Teams invent elaborate DSLs for this and then nobody writes scenarios. The ones that get written fit on one screen.
Deriving scenarios from the spec
The OpenAPI document already tells you where journeys begin and end:
- Find the state machines. Any schema with a
statusenum is a scenario waiting to happen. List every legal transition (pending → paid → fulfilled,draft → published → archived) and write one happy-path scenario plus one illegal-transition scenario. - Follow resource nesting.
POST /orders→GET /orders/{id}→POST /orders/{id}/itemsis a parent-child chain. Theparametersblocks literally name the ids that must be extracted. - Pair every write with its read. After a
PATCH, GET the resource and assert the patched fields. This catches the bug where the update endpoint accepts fields it silently discards. - Use
4xxresponses as test cases. The spec documents a409on duplicate creation — write the scenario that creates twice and expects it. Error branches are where contract documentation is usually thinnest, so scenarios double as spec verification. - Cover pagination and idempotency. A scenario that creates three resources and pages through the list verifies cursor behavior; replaying a create with the same idempotency key verifies the
Idempotency-Keycontract.
Polling and SSE: time is part of the test
Modern APIs are rarely request-response all the way through. Two patterns need first-class handling:
Polling. A job endpoint returns 202 with a status URL. The scenario needs a poll step: GET until status == ready or a timeout, with a declared interval. Asserting on the first response is the classic flaky test — on a fast machine the job is sometimes already done, on CI never.
Server-sent events. When the spec describes a stream (text/event-stream with an x-protocol extension carrying event names and item schema), the scenario opens the connection, triggers the producing action, and asserts that the expected event arrives within a window. For example: open the project event stream, update a setting, expect one project.updated event with the new value, then close. Testing this with plain HTTP clients means hand-rolling an event parser per service; a spec-aware runner already knows the event contract.
Two environments, same scenario
The highest-leverage habit is running the identical scenario against two targets:
- The mock, before the backend exists. Frontend and backend agree on the journey by executing it against generated responses. Missing fields become arguments in week one instead of bugs in week six.
- Staging, after it exists. The scenario file does not change; only the server URL and credentials do. When the journey passes on mock but fails on staging, the implementation diverged from the contract — that is the single most valuable signal a contract workflow produces.
In CI, gate the merge on the mock run and run the staging run on a schedule or before release, with real (seeded) accounts.
What a failing scenario report should say
"Assertion failed" is not enough. A useful report names the step, shows expected versus actual at the exact JSON path, includes the full exchange (request, response headers, body), and — critically — prints the extracted variables so you can see that order_id came back undefined before the second request 404'd. Most of the time the bug is an extraction path that drifted when a wrapper envelope was added.
Start with three scenarios
Do not attempt coverage on day one. Write the three journeys that map to the three things customers pay to do. For a billing API that is: subscribe, change plan at period end, cancel. For a document service: upload, process (poll), download. Run them against the mock Monday, against staging Friday. Everything else is incremental.
The desktop workspace runs these scenarios against the same OpenAPI document used for design and mocking, including the polling and SSE steps — the online demo includes a sample spec with a streaming operation. The design-first angle is covered in designing an API before the first line of code.
What to read next: A 200 is not done — run the business scenario is the short version of this argument, and from a scanned spec to mocks, scenario tests, and an MCP server shows the flow starting from legacy code.