One MCP Gateway for All Your Internal APIs: The Aggregation Pattern
A single MCP server in front of one service is a solved problem: generate tools from the OpenAPI spec, run it over stdio or HTTP, done. At company scale the problem changes shape. A midsize platform team has forty services, each with its own spec, its own auth, its own staging and production hosts. Let every team publish an MCP endpoint and you quickly get:
- Agents configured against dozens of URLs, each with its own OAuth consent.
- Tool catalogs in the hundreds, past the limit most clients expose to the model, so tools silently disappear.
- Name collisions: three
list_users, twocreate_order, no way to tell them apart. - No central point for rate limiting, audit logs, or revocation.
The answer that keeps working as services multiply is an MCP gateway: one endpoint an agent authenticates against, which aggregates many backend APIs into one namespaced, governed tool catalog.
The shape of the pattern
AI clients (Claude Desktop, Cursor, VS Code, CI agents)
│ OAuth 2.1, one consent, one token
▼
┌──────────────── MCP gateway ────────────────┐
│ auth & scopes rate limits audit log │
│ catalog aggregation name normalization │
└───────┬───────────┬───────────┬──────────────┘
▼ ▼ ▼
orders API billing API support API
(OpenAPI) (OpenAPI) (OpenAPI)
│ │ │
▼ ▼ ▼
services + databases (private network)The gateway is not a new API implementation. It is a composition layer: each backend keeps owning its spec and its service; the gateway compiles those specs into one MCP catalog and proxies calls.
Catalog aggregation: namespace, do not flatten
The core design decision is how services appear in tools/list. Flattening produces collisions and ambiguity. Prefixing every tool with a service namespace produces a predictable, greppable catalog:
{
"tools": [
{
"name": "orders__create_order",
"description": "[orders] Create a pending order and reserve inventory for 15 minutes.",
"inputSchema": { "$ref": "#/$defs/orders.CreateOrderRequest" }
},
{
"name": "billing__create_invoice",
"description": "[billing] Issue an invoice for a fulfilled order.",
"inputSchema": { "$ref": "#/$defs/billing.CreateInvoiceRequest" }
}
]
}Three rules keep the catalog usable:
- Namespaces are stable and short (
orders,billing), taken from a registry rather than guessed from repo names. - Schema components are namespaced too (
orders.Order,billing.Invoice) so$refresolution never merges two teams'Errormodels. - Descriptions carry the namespace as a prefix, because models read descriptions more reliably than tool-name conventions.
The gateway should also deduplicate genuinely shared schemas by reference rather than copying them, so the model sees one Money concept, not twelve.
Tool budget: more services does not mean more exposed tools
Clients and models have practical limits on tool count; hundreds of tools degrade selection accuracy even where the client technically accepts them. A gateway earns its keep by managing that budget:
- Capability scopes filter the catalog. A token with
orders:read,billing:readsees only those namespaces;tools/listitself is authorization-aware. - Task-scoped views expose a curated subset for common workflows ("incident triage", "order-to-cash") instead of the entire platform.
- Read and write tools split by scope, so broad onboarding can start read-only; the same principle as MCP least-privilege design.
- Deep query operations stay resources. Search endpoints and document lookups fit MCP resources better than dozens of getter tools; see tools vs resources vs prompts.
Auth: one token in, service credentials out
The agent authenticates once against the gateway using OAuth 2.1 with PKCE. The gateway then holds service-to-service credentials downstream and maps the caller's identity and scopes onto each request:
agent token (scopes: orders:read, billing:write)
│
▼
gateway authorizes orders__get_order → allowed, signs request as gateway, forwards user identity
gateway authorizes billing__create_invoice → allowed, mints scoped service token
gateway authorizes admin__delete_account → denied at the edge with a structured MCP errorForward the authenticated user's identity to backends in a signed header or token exchange rather than making every call anonymous; otherwise audit trails end at the gateway. The OAuth mechanics at the edge are the same ones described in MCP authentication explained.
The operational features that belong at the edge
Because every tool call crosses the gateway, concerns that would be duplicated forty times are implemented once:
- Rate limiting and quotas, per user and per namespace, with
Retry-Afteron throttles. - Audit logging in one schema, including the calling client, scope used, and arguments with PII redacted.
- Versioning and deprecation: the gateway can serve
orders__create_orderandorders__create_order_v2side by side while backends migrate, following the approach in versioning MCP tools. - Observability: latency and error-rate metrics per namespace give teams feedback on how agents actually use their APIs.
- Fail containment: a backend outage degrades one namespace and returns a clean, retryable MCP error instead of failing the whole catalog.
Build vs buy, and what to do on day one
Do not start by building a platform. The aggregation pattern can be introduced incrementally:
- Week one: publish one generated MCP server for the highest-value service, remotely, with OAuth and logs.
- Week two: put a thin gateway in front of it that does nothing but auth and pass-through, so clients already point at the stable URL.
- As demand grows: register additional services by adding their OpenAPI specs to the gateway's catalog build, each in its namespace, without clients changing their config.
- Only when needed: add scoped catalog views, quotas, and curated task bundles.
Because the tool catalog is generated from specs, adding a service is a configuration and build step — register the spec, declare the namespace, choose the scopes — rather than an integration project. Treating the MCP server as compiled output is the principle in the MCP server is a build artifact.
Where the specs come from
Teams adopting this pattern usually discover that the prerequisite is not MCP infrastructure at all; it is having trustworthy, current OpenAPI documents for every service. For services that predate spec-driven development, those can be generated from the running codebase and then maintained as the contract; the code-to-spec workflow is covered in generating OpenAPI from existing code.
A local-first API workspace can generate the individual servers and validate the aggregated catalog before anything is hosted; you can try generating a server from a spec in the online demo, and the client-side configuration for the gateway URL is compared across editors in MCP clients compared.