# How to organize a large OpenAPI spec: multi-file structure, $ref rules, and CI checks

Every OpenAPI document starts as one clean file. Then it reaches forty operations and the file is 2,000 lines, then a hundred operations and pull requests conflict on every merge, then nobody edits it without a full afternoon of context. The standard advice is "split it with `$ref`," which is true and unhelpful — split *how*, along which boundaries, and what breaks in tooling when you do? This is the layout and rule set that survives 200+ operation APIs.

## The target structure

Organize by artifact kind first, domain second. Paths and schemas are different rates of change and different authors touch them, so they do not share folders.

``` text
openapi/
  openapi.yaml              # entry point: info, servers, security, tags
  paths/
    projects/
      projects.yaml         # /projects collection
      project-item.yaml     # /projects/{projectId}
      project-history.yaml  # /projects/{projectId}/history
    billing/
      subscriptions.yaml
      invoices.yaml
  schemas/
    common/
      pagination.yaml
      errors.yaml
      money.yaml
    project/
      project.yaml
      project-create.yaml
      project-event.yaml
    billing/
      subscription.yaml
      invoice.yaml
  responses/
    errors.yaml             # shared 400/401/403/404/409/429/500
  parameters/
    pagination.yaml
    ids.yaml
  examples/
    projects/
      project-example.json
  callbacks/
    webhooks.yaml
  x-events/
    project-events.yaml     # SSE event contracts
```

The entry document stays small — it exists to assemble, not to describe:

``` yaml
openapi: 3.2.0
info:
  title: Platform API
  version: "2.4.0"
servers:
  - url: https://api.example.com/v1
tags:
  - name: projects
  - name: billing
paths:
  /projects:
    $ref: paths/projects/projects.yaml#/~1projects
  /projects/{projectId}:
    $ref: paths/projects/project-item.yaml#/~1projects~1{projectId}
components:
  schemas:
    Project:
      $ref: schemas/project/project.yaml#/Project
    ProjectCreate:
      $ref: schemas/project/project-create.yaml#/ProjectCreate
  parameters:
    PageCursor:
      $ref: parameters/pagination.yaml#/PageCursor
  responses:
    BadRequest:
      $ref: responses/errors.yaml#/BadRequest
```

Two practical notes. First, some teams prefer each path file export a full Path Item under a simple anchor rather than URL-encoded JSON-pointer keys (`~1` for `/`); a bundler normalizes either, so pick one convention and lint for it. Second, keep the entry point as the only file tooling needs to find — everything resolves from it, never from a folder scan.

## The $ref conventions that matter

Most multi-file pain is self-inflicted through inconsistent referencing. Five rules cover almost all of it:

1.  **One canonical definition per artifact.** A schema lives in exactly one file; everything else references it. Duplicating a `Money` type into two folders "to avoid a dependency" guarantees divergence.
2.  **Reference component objects, not inline fragments.** Point at `#/Project`, never at a sub-property of someone else's schema. If a sub-object is reused, promote it to its own named schema.
3.  **Paths reference components; components never reference paths.** Dependencies point inward. Schemas referencing path files create cycles that bundlers resolve inconsistently.
4.  **Prefer same-file refs for tightly coupled private types.** A `ProjectStatus` enum used only by `Project` stays in `project.yaml`. Splitting every enum into its own file produces a folder of 150 three-line files nobody navigates.
5.  **No remote refs to repos you do not pin.** `$ref: https://example.com/schemas/common.yaml#/X` makes your build depend on someone else's uptime and makes builds non-reproducible. Vendor shared schemas into the repo and update them deliberately.

## Bundling and resolution

Source layout is multi-file; published artifacts are usually bundled. The pipeline:

-   **Author and review** in the split structure — diffs are small and ownership is clear (CODEOWNERS on `paths/billing/`).
-   **Bundle for publication** into a single dereferable document for docs renderers and gateways that do not follow relative refs.
-   **Keep refs in generated SDKs** where the generator supports them, so consumers get named types (`Project`) rather than 400 anonymous inline interfaces.

Validate with two resolver modes in CI: strict (any unresolvable ref fails the build) and bundled (the final single-file output must itself validate, because bundling can mask or duplicate anchors).

## Common traps

-   **Circular refs across files.** Project references Owner; Owner references a list of Projects. Circular refs are legal within components but some renderers and older generators stack-overflow on them. Test the toolchain early; when in doubt, break the cycle with a summary type (`OwnerRef: { id, displayName }`).
-   **`$ref` siblings.** In 3.0 a `$ref` ignored its siblings, which made overriding descriptions impossible. 3.1+ allows sibling keywords, but verify your renderer and generator honor them rather than silently dropping the description.
-   **Merge artifacts (allOf) used as import systems.** `allOf` is for composition, not for pulling fields across files because the folder structure was inconvenient. If you find five layers of allOf to assemble one request body, flatten the domain model.
-   **Examples drifting from schemas.** Keep examples as separate JSON files referenced from media-type objects, and validate example payloads against their schemas in CI — an example that fails its own schema is worse than none.
-   **Generated folders committed by mistake.** `dist/` bundles next to sources invite edits to the generated file. Keep build output out of the source tree.

## The CI gate

A spec this size needs automated discipline; a human review at 2 a.m. will not provide it:

1.  Structural validation against OpenAPI 3.2.
2.  Reference resolution with no external-network access (catches remote-ref drift).
3.  Linting against a house style guide: operationId naming, every operation has summary + description + at least one example, tags exist in the root list, no undocumented error responses.
4.  Example validation: every example validates against its schema.
5.  Breaking-change diff against the released version (see [detecting breaking changes in CI](https://www.powerduck.com/blog/detect-breaking-api-changes-openapi-diff-ci/)).
6.  Bundled-output validation, because the artifact consumers load is the bundle.

## When not to split

Multi-file is not free. Under \~30 operations, a single well-organized file is faster to search and every tool handles it without configuration. Split when you have multiple regular editors, code ownership boundaries, or merge conflicts on the spec — not before. And if you inherit a monolith, split incrementally: move `components.schemas` to files first (lowest risk, pure renames), then paths by domain, validating the bundle after every move.

Powerduck opens the split folder as a workspace document, resolves the graph locally, and keeps mocks, scenario tests, docs, and MCP serving bound to the same source tree — so the multi-file layout is what gets authored and what gets run, with no separate bundling step to maintain by hand. The [quickstart](https://www.powerduck.com/docs/overview/quickstart?ref=powerduck.com) walks through opening a structured document.

**What to read next:** [Swagger 2.0 to OpenAPI 3.x migration](https://www.powerduck.com/blog/swagger-2-to-openapi-3-migration/) often precedes a reorganization like this, and [one spec as the source of truth](https://www.powerduck.com/blog/one-spec-source-of-truth/) explains why the source tree should feed every downstream artifact.

