Generate OpenAPI 3.2 from existing code: Express, FastAPI, Go, and Spring compared
Teams with no spec usually reach for one of three recovery methods: bolt annotations onto every handler, put a proxy in front of staging and record traffic, or scan the source. All three produce an OpenAPI document; only one of them captures routes that nobody has called recently, and only two of them capture schemas without hand-writing anything. Here is how the approaches compare, what the AST-based scanner in Powerduck extracts per framework, and where each method lies to you.
The three approaches
Annotations and framework-integrated generation. FastAPI, NestJS with decorators, Spring with springdoc, and ASP.NET with Swashbuckle generate specs from metadata in the code. Accuracy is excellent if the annotations are complete, which they never are on a legacy service — the whole reason you have no spec. Retrofitting annotations across 200 routes is a multi-month project, and untyped dynamic handlers stay undescribed no matter how many decorators you add.
Traffic recording. A proxy observes real requests and infers endpoints and shapes. Setup is fast and the data is undeniably real — but it only contains exercised paths. Error branches, admin routes, the endpoint for the customer who churned, and anything behind a feature flag are invisible forever. Recorded payloads also leak production data into the spec, and inferred nullability is wrong systematically: a field absent in one recorded response is not the same as an optional field.
AST scanning. A parser builds the syntax tree, follows router registration across files, resolves types into schemas, and emits OpenAPI. It sees every route, including dead ones, runs entirely on your machine, and repeats cheaply after every sprint. Its weakness is the mirror image of traffic recording: it can only describe what the code structurally contains, so genuinely dynamic handlers produce honest gaps instead of schemas.
The strongest workflow combines them: scan for completeness, record traffic to enrich examples, and annotate the handful of genuinely ambiguous handlers.
What the scanner extracts, by stack
The scanner walks framework-specific routing on top of language grammars. Coverage as of the 0.9 line:
| Language | Frameworks | What resolves into schemas |
|---|---|---|
| TypeScript / JavaScript | Express, Fastify, NestJS, Koa, Hono, Elysia, Next.js route handlers | TS interfaces/types, Zod, ArkType, TypeBox contracts, fluent-json-schema, DTO classes |
| Python | FastAPI, Flask, Django REST Framework | Pydantic models, dataclasses, typed handlers; Flask view models where present |
| Go | Gin, Chi, Echo, net/http, gorilla/mux | structs with json tags, ShouldBindJSON, json.NewDecoder, pointer/value bodies |
| Java | Spring MVC/WebFlux, JAX-RS (Jersey) | @RequestBody/@Valid DTOs, records, response entities, bean validation |
| C# | ASP.NET Core (controllers and minimal APIs) | model classes, data annotations, [FromBody]/[FromQuery] |
| Rust | Axum, Actix-Web | serde structs, extractors, Json<T> bodies |
| PHP | Laravel, Symfony | form requests, resource classes, typed DTO patterns |
HTTP and SSE are both first-class outputs; gRPC and GraphQL stay labeled rather than being forced into REST shapes.
Strongly typed vs loosely typed handlers
The extraction strategy deliberately differs by language, because the evidence available differs.
In a Go Gin handler, the contract is largely in the types. c.ShouldBindJSON(&req) where req is CreateProjectRequest{ Name string \json:"name" binding:"required"\; Plan *Plan \json:"plan"\ } yields a request body schema with a required name and a nullable plan, with no model call involved. Java DTOs with Jakarta validation annotations do the same, including @Size, @Pattern, and @Email. These operations are extracted deterministically; an AI step would only add cost and variance.
In a legacy Express handler the evidence is thinner: app.post("/projects", (req, res) => service.create(req.body)). The scanner knows the route, the method, the path parameters, and often the response shape from what the handler returns, but the request body is untyped. Two choices exist, and the product takes the boring one: mark body-schema-unknown as an explicit gap in the review dialog, or — if the user enables AI enhancement — send only that handler's source to the model configured in the workspace, get a proposed schema, and require one-click approval before it enters the document. The model never sees the whole codebase, and nothing is imported silently.
The routing problems that text search fails
Real codebases defeat grep-based extraction in predictable ways, and the AST walk handles each:
- Mounted routers with prefixes.
app.use("/v1/admin", adminRouter)→router.use("/projects", projectRouter)→projectRouter.post("/:id/history/compare", ...). The full path is reconstructed across three files; naïve regex produces/:id/history/compare. - Cross-file route contracts. Hono's
createRoutein one file, the handler in another; Elysia's third-argument options object with chained.get/.groupprefixes. - Indirect binding. Echo handlers that bind inside a helper (
req.bind(c, &u)), Go split-statement decoders (dec := json.NewDecoder(r.Body); dec.Decode(&u)),new(T)bodies. - Multiple response DTOs per status. A controller returning
ResponseEntity<ProjectDto>on success andErrorResponseon failure becomes two explicit responses, not one optimistic 200. - draft-07 schemas. Fastify
definitions/$defsare hoisted intocomponents.schemaswith references rewritten.
The review gate and honest gaps
Nothing imports without a review. The scan dialog lists every operation with method, full path, confidence, and a gap report: body-schema-unknown, response-schema-unknown, auth-unknown, sse-events-unknown. A field marked unknown in the finished document was first shown to the user as a gap. This is the opposite failure mode from traffic tools that confidently emit "field1": "string" — a flagged gap can be fixed deliberately; a confident guess fails in production.
Rescans are diffs
After the first confirmed import, a .powerduck/discovery.json sidecar records fingerprints. The next scan produces a change set — added routes, removed routes, renamed parameters — and merges new operations while preserving descriptions, examples, and manual edits. That is what makes the document survive past the initial enthusiasm: keeping it honest costs a rescan and a five-minute review instead of another reverse-engineering project.
The scanner is the open-source @powerduck/code-to-openapi engine; the inherited-service war story shows it running across Express and FastAPI services with ~200 routes.
What to read next: from a scanned spec to mocks, scenario tests, and an MCP server covers day two after the scan, and a curl command is not an API handoff covers the lighter-weight import path.