Every production API rate-limits. Most document it as a single sentence: "100 requests per minute." The client never learns its remaining budget, never knows when the window resets, and treats a 429 the same as a 500. Autonomous agents are worse: they retry immediately and get banned. A rate-limit contract is not the limit itself; it is the machine-readable feedback that lets a well-behaved client stay under it.

Pick the right status code

CodeMeaningTypical cause
429 Too Many RequestsThe client is over its quotaPer-token, per-user, or per-IP rate limiting
403 ForbiddenIdentity is known but not allowedPlan does not include the endpoint, not a rate issue
503 Service UnavailableThe whole service is degradedOverload or maintenance; pair with Retry-After

Do not return 403 for throttling and do not return 429 for a plan restriction. 429 specifically means "too many requests in a window," and it is the one response that must carry retry guidance.

Retry-After is mandatory on a 429

The header takes one of two forms.

Delta seconds, the common case for a fixed window:

HTTP/1.1 429 Too Many Requests
Retry-After: 42

An HTTP date, used when the block lasts until an absolute instant:

Retry-After: Wed, 07 Oct 2026 15:00:00 GMT

Clients and agents implement the delta form trivially (sleep, then retry once). A date needs clock handling, so prefer seconds for ordinary throttling. Never emit Retry-After: 0 on a 429; it tells the client to slam the endpoint again immediately and defeats the limit.

Report the budget, not just the block

The RateLimit-* headers from the IETF rate-limit header draft give the client a running picture on every response, including successful ones.

RateLimit-Limit: 100; window=60
RateLimit-Remaining: 37
RateLimit-Reset: 23
  • RateLimit-Limit is the quota and the window in seconds.
  • RateLimit-Remaining counts down to zero.
  • RateLimit-Reset is seconds until the window refreshes.

Many gateways still emit the older X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset names. Document whichever set your gateway actually sends; do not list headers your edge layer does not produce. If you control the gateway, prefer the unprefixed draft names.

Model it once, reuse it everywhere

Define the shared headers and a problem body, then reference the 429 response from every operation.

components:
  headers:
    RateLimitLimit:
      description: Requests allowed per window for this credential.
      schema: { type: integer, example: 100 }
    RateLimitRemaining:
      description: Requests left in the current window.
      schema: { type: integer, example: 0 }
    RateLimitReset:
      description: Seconds until the window resets.
      schema: { type: integer, example: 23 }
    RetryAfter:
      description: Minimum seconds to wait before retrying.
      schema: { type: integer, example: 42 }
  responses:
    TooManyRequests:
      description: Quota exceeded. Honor Retry-After with exponential backoff and jitter.
      headers:
        Retry-After: { $ref: '#/components/headers/RetryAfter' }
        RateLimit-Limit: { $ref: '#/components/headers/RateLimitLimit' }
        RateLimit-Remaining: { $ref: '#/components/headers/RateLimitRemaining' }
        RateLimit-Reset: { $ref: '#/components/headers/RateLimitReset' }
      content:
        application/problem+json:
          schema:
            allOf:
              - $ref: '#/components/schemas/Problem'
              - type: object
                properties:
                  retryAfterSeconds: { type: integer, example: 42 }
                  limit: { type: string, example: '100; window=60' }
          example:
            type: https://api.powerduck.com/probs/rate-limited
            title: Too Many Requests
            status: 429
            detail: 'Plan free allows 100 requests per 60 seconds.'
            retryAfterSeconds: 42

Then each operation gets the response for free:

responses:
  '200': { $ref: '#/components/responses/InvoiceList' }
  '429': { $ref: '#/components/responses/TooManyRequests' }

State the quota dimension and scope

"100 per minute" is incomplete. Say what the counter keys on, because clients share credentials and run background jobs.

DimensionExampleNote
Credential100 req/min per API keyMost common; document shared-account behavior
User60 req/min per userRelevant when one key serves many users
IP300 req/min per IPCatches anonymous traffic; flag NAT false positives
Endpoint class10 expensive exports/minHeavy routes often have their own bucket
Concurrency2 concurrent exportsA separate limit; return 429 or 503 consistently

Put the dimension, window, and per-plan differences in the operation description or a dedicated rate-limit section. A free and a paid tier with the same 429 shape but different limits should show both values.

Verify the backoff before production

A documented 429 is only useful if the client handles it. Drive a mock server from the spec and script the edge cases:

# First call exhausts the window, second returns the throttle contract.
curl -i https://mock.local/invoices \
  | grep -iE 'ratelimit|retry-after'

curl -i https://mock.local/invoices
# HTTP/1.1 429
# Retry-After: 42
# RateLimit-Remaining: 0

Assert three behaviors: the client reads Retry-After instead of using a fixed sleep, it adds jitter so a fleet of throttled agents does not retry in lockstep, and it surfaces a clear error after the window still fails instead of looping forever. A spec-driven mock can return the 429 on demand, which is far easier than flooding a real gateway.

For AI agents, Retry-After is the single most important header in the contract: it turns an unstructured refusal into a scheduling instruction. Agents that parse it pause exactly as long as asked and resume without human intervention.

Common mistakes

  • 429 with no Retry-After, forcing clients to guess.
  • Headers documented in the spec but never emitted by the gateway (always capture real response headers and align the two).
  • Using 503 for per-user throttling, which conflates one noisy client with an outage.
  • Resetting a fixed window such that a client sends 100 requests at 00:59 and 100 more at 01:00. A sliding window or a documented fixed boundary fixes the burst; say which one you run.
  • Hiding plan limits. A client on a paid tier should see its higher quota in RateLimit-Limit, not discover it by being blocked.

Checklist

  1. Throttling returns 429, never 403; overload returns 503.
  2. Every 429 carries a non-zero Retry-After in seconds.
  3. Successful responses carry the RateLimit-* budget headers your gateway actually sends.
  4. A reusable 429 response with a problem body is referenced by every operation.
  5. The spec states the quota dimension, window, and per-plan values.
  6. A mock reproduces the 429 and the client is tested for backoff, jitter, and a final give-up.

When the limit is visible before the block, well-behaved clients never hit the wall.

Describe the throttled response, run it through a spec-driven mock, and confirm your client backs off correctly, all in the browser app. The error body above follows the standard problem-details format; see the RFC 9457 error response guide for the full envelope.