Recursive and cyclic schemas in OpenAPI: self-referencing $ref for trees, nested comments, and graphs
Every content API eventually meets a shape that contains itself: a category tree, a comment thread, a folder of folders, an org chart, a bill of materials. The JSON arrives nested five or fifty levels deep, and the first OpenAPI attempt either inlines the object five times and gives up, or types children as a generic array of object and loses every field.
The fix is a schema that references itself. It is one of the smallest and most misunderstood mechanisms in JSON Schema.
A tree is a self-referencing schema
A comment has replies, and a reply is a comment. Define the object once in components.schemas and point a property back at it:
components:
schemas:
Comment:
type: object
required: [id, body, created_at]
properties:
id:
type: string
format: uuid
body:
type: string
author:
$ref: '#/components/schemas/Author'
created_at:
type: string
format: date-time
replies:
type: array
items:
$ref: '#/components/schemas/Comment'That single $ref is the whole trick. There is no depth keyword and no need to declare recursion. A schema is allowed to reference itself, directly or through another schema, and a compliant resolver follows the loop lazily rather than expanding it.
A TypeScript generator emits exactly the recursive interface you would write by hand:
export interface Comment {
id: string;
body: string;
author?: Author;
created_at: string;
replies?: Comment[];
}Python generators emit a forward-referenced model ("Comment" in quotes, resolved later), Go generators emit a struct with a slice of pointers to the same struct, and Java generators emit a self-referential bean. All of them rely on the single fact that the type is named and reused; an anonymous inline object has no name to recurse through.
Empty array, omitted, or null: pick one for leaves
A leaf node is where most recursive specs get vague. The three encodings below mean different things and clients branch differently:
| Encoding of a leaf | Meaning | Client check |
|---|---|---|
"replies": [] | There are no replies, and the field is always present | replies.length === 0 |
replies omitted | Replies were not loaded or do not apply | 'replies' in comment |
"replies": null | Explicitly known to have no replies (3.1 nullable) | replies === null |
Choose one and document it. The cheapest contract for a tree is "always present, empty at the leaf." It removes the null/omitted ambiguity and makes pagination of children straightforward. If replies are loaded lazily, omitting the property until the next request is more honest than sending an empty array, because empty then unambiguously means "no children."
Recursive alternatives with oneOf
A filesystem node is either a file or a folder, and only the folder has children. Recurse through a discriminated union instead of one object with every field optional:
Node:
type: object
discriminator:
propertyName: kind
properties:
kind:
type: string
name:
type: string
oneOf:
- $ref: '#/components/schemas/FileNode'
- $ref: '#/components/schemas/FolderNode'
FolderNode:
allOf:
- $ref: '#/components/schemas/Node'
- type: object
properties:
kind:
type: string
const: folder
children:
type: array
items:
$ref: '#/components/schemas/Node'The recursion now goes FolderNode -> Node -> (FileNode | FolderNode), which is mutually recursive across three schemas. Generators that support discriminators produce a tagged union and a compiler proves that a file never carries children. If you instead put children on the base node, every file gets a meaningless optional array.
Trees nest; graphs point
Real recursion only works for data that is a tree, where each node has one parent and can be physically nested. A graph (followers, related products, a directed acyclic workflow) has nodes that point at each other, and JSON cannot nest a cycle without infinite expansion. Graphs use references by identity:
Graph:
type: object
properties:
nodes:
type: array
items:
$ref: '#/components/schemas/GraphNode'
edges:
type: array
items:
$ref: '#/components/schemas/GraphEdge'
GraphNode:
type: object
required: [id]
properties:
id: { type: string }
label: { type: string }
GraphEdge:
type: object
required: [from, to]
properties:
from: { type: string, description: References GraphNode.id }
to: { type: string, description: References GraphNode.id }Do not be tempted to model a graph as nested objects that "$ref back." JSON references only exist inside the schema; the payload itself has no pointers, so a cyclic object graph must be flattened into nodes and edges (or nodes with an included map, the JSON:API approach). The schema can express "this string is the id of a node in the same document" only in a description; JSON Schema has no native foreign-key keyword.
Mutual recursion and $ref siblings
Schemas can recurse through each other: Book -> Chapter -> Section -> Section. The rule is the same, every participant must be a named component schema, because only named schemas have a fragment identifier to point at. Two practical notes:
- In OpenAPI 3.1,
$refis a plain JSON Schema keyword and can sit alongside annotations ($refplusdescriptionornullableviatype). In 3.0, siblings next to$refwere ignored, so wrap withallOfif you are still on 3.0. - Keep the cycle inside
components.schemas. A recursive schema defined inline under a path has no stable identifier to reference.
What static scanners and mocks need
When an OpenAPI document is reverse-engineered from code, recursion is a common place a naive parser flattens or gives up. A sound parser follows named type aliases and ORM relations, detects when it reaches a type already on the current path, and emits a $ref back to that type instead of expanding forever. It should distinguish an array of children (recurse) from a foreign-key id (a graph edge), and it should report an unknown rather than silently capping the tree at depth two.
A spec-driven mock that understands recursion can generate a tree of controlled depth (two or three levels, empty arrays at the leaves) so the UI has realistic nested data without an infinite payload.
Checklist
- Define the recursive object once under
components.schemas; point children back with$ref, never inline-and-repeat. - Decide the leaf encoding up front: empty array, omitted, or nullable, and use it everywhere.
- Use a discriminated
oneOfwhen only some variants recurse (file versus folder). - Flatten true graphs into nodes and edges (or nodes plus an included map); do not try to nest cycles.
- Make every participant of mutual recursion a named schema.
- On 3.0, wrap
$refsiblings inallOf; on 3.1 you can annotate alongside$ref. - Generate the client and confirm you get a recursive type that compiles, and generate a mock tree that terminates.
Get these right and a comment thread of any depth is the same amount of schema as a single comment.
You can build recursive and mutually recursive schemas, generate a client whose types compile, and mock a nested tree that terminates, in one local-first workspace, right in your browser. For the rules that keep a component library from turning into reference spaghetti, see reusable JSON Schema components in OpenAPI.