Cursor vs offset pagination: what to actually put in your OpenAPI spec
Pagination is one of those decisions that gets made in twenty minutes and paid back over the entire life of an API. Every list endpoint has it, every SDK generates it, and every frontend engineer has a story about a duplicate row that appeared on page two — or a row that vanished between page one and page two. The cursor-versus-offset debate is old, but the right answer is concrete once you look at how the data actually changes and what the OpenAPI document needs to say to make clients reliable.
Offset pagination: simple, stable only on frozen data
Offset (or page-number) pagination takes ?page=2&per_page=50 or ?offset=50&limit=50 and translates it into LIMIT/OFFSET or its equivalent.
It has two genuine advantages: clients can jump to any page, which makes traditional paginated UIs with page numbers possible, and the parameters are self-explanatory — no opaque tokens to debug.
The failure mode is drift as the underlying set changes. Sort by created_at desc; a new row is inserted while the user is on page one; page two now repeats the last row of page one, because everything shifted by one. Delete a row and the opposite happens: a row is skipped. On top of that, deep offsets are a database performance problem — OFFSET 100000 still scans and discards 100,000 rows in most engines.
Documented in OpenAPI:
parameters:
- name: page
in: query
schema:
type: integer
minimum: 1
default: 1
- name: per_page
in: query
schema:
type: integer
minimum: 1
maximum: 100
default: 20Use it when: the dataset is small and slow-moving (settings lists, admin screens), clients genuinely need random page access, and consistency across pages is not a correctness requirement.
Cursor pagination: stable for changing data
Cursor pagination returns an opaque token representing a position in the ordered set:
GET /v1/events?limit=50
HTTP/1.1 200 OK
{
"data": [ ... ],
"page_info": {
"next_cursor": "eyJ0cyI6MTcyMzQ1Njc4OX0",
"has_more": true
}
}The cursor encodes the sort key of the last item (often a composite of timestamp and id, base64-encoded). The next query is ?cursor=eyJ0cyI6...&limit=50. Because the position is anchored to a value rather than a row count, inserts and deletes do not shift the window: new rows appear on the first page without duplicating rows on subsequent pages, and indexes make the query equally fast at depth 100,000. This is why Stripe, GitHub, and Slack standardized on it for activity and event streams.
The costs are real: no jump-to-page, clients cannot derive a total page count cheaply (a total_count is an expensive and often stale number on large sets), and cursors must be treated as opaque — clients that decode and construct them break the moment you change encoding.
Documenting it correctly in OpenAPI means saying all of that:
parameters:
- name: cursor
in: query
description: Opaque pagination token from the previous response's page_info.next_cursor. Do not parse or construct.
schema:
type: string
- name: limit
in: query
schema:
type: integer
minimum: 1
maximum: 100
default: 50Keyset is the honest middle child
A third option is keyset pagination, where the client passes the last seen key directly: ?since_id=10432. It is stable and index-friendly like cursors, but it is not opaque, so clients can reason about it — and construct it, which means you can never change the sort key without a breaking change. Keyset works well for internal APIs with one immutable ordering; cursors are keyset wrapped in an abstraction that buys you the freedom to evolve.
How to choose
| Question | Offset | Cursor | Keyset |
|---|---|---|---|
| Data changes while paginating? | Duplicates/skips | Stable | Stable |
| Need jump-to-page UI? | Yes | No | No |
| Deep pages performance-sensitive? | Bad | Constant | Constant |
| Sort order may evolve? | Fine | Token hides it | Breaking |
| Need exact total count? | Cheap | Often dropped | Often dropped |
| Typical use | Admin tables, small lists | Feeds, events, logs, search | Internal fixed-order lists |
A common and defensible setup: cursor pagination on activity/event/resource-list endpoints where clients walk forward, and offset retained on a handful of back-office tables where page numbers are the UI.
The response envelope belongs in components
Whatever the style, define the envelope once and reuse it, because inconsistent list responses are the bug generator in practice:
components:
schemas:
CursorPage:
type: object
description: Generic cursor-paginated result wrapper.
required: [data, page_info]
properties:
data:
type: array
items: {}
page_info:
type: object
required: [has_more]
properties:
next_cursor:
type: [string, "null"]
description: Present only when has_more is true.
has_more:
type: boolean
ProjectCursorPage:
allOf:
- $ref: '#/components/schemas/CursorPage'
- type: object
properties:
data:
type: array
items:
$ref: '#/components/schemas/Project'Rules that save support tickets: always include has_more (clients should not infer the end from a short page — the last page can legitimately be full); return next_cursor: null rather than omitting it, so generated clients type it consistently; and document whether the cursor encodes a snapshot (some implementations pin the client to a point-in-time view — say so explicitly, it is a feature).
Testing the contract
Pagination is where scenario tests earn their keep. A single request cannot catch the duplicate-row bug; a journey can: seed 65 records, walk all pages with limit=20, assert every id appears exactly once, assert termination when has_more is false, then insert a record mid-walk on a cursor endpoint and assert no duplication. Run the same scenario against the mock and staging and the contract — not just the happy path — is verified. That chain is exactly what spec-driven scenario runners automate from the OpenAPI document.
Powerduck runs those pagination journeys against local mocks and staging from the same spec that renders the docs, and the workspace validates envelope schemas while you design the endpoint. The demo includes a paginated collection to experiment with.
What to read next: scenario testing for REST APIs shows the full create-page-walk assertions, and REST error responses with RFC 9457 covers the other envelope every list endpoint needs.