# File downloads in OpenAPI: CSV, PDF, ZIP, images, Content-Disposition, and non-JSON responses

Most OpenAPI tutorials stop at `application/json`, yet a large share of real endpoints hand back a file: an invoice PDF, a CSV export, a ZIP of assets, a resized image, a generated spreadsheet. Teams either leave these endpoints undocumented or describe them as `type: string` with no format, so a generated client tries to decode a PDF as UTF-8 text and corrupts it. Non-JSON responses need explicit media types, a binary schema, and the headers that tell the client what to do with the bytes.

## The schema for a binary body

A file response uses `type: string` with `format: binary` (OpenAPI 3.x), not an object and not an untyped string. The media type says what the bytes are:

``` yaml
paths:
  /invoices/{invoiceId}/pdf:
    get:
      summary: Download an invoice as PDF
      operationId: downloadInvoicePdf
      parameters:
        - name: invoiceId
          in: path
          required: true
          schema: { type: string }
      responses:
        '200':
          description: The invoice PDF.
          headers:
            Content-Disposition:
              schema: { type: string }
              description: 'attachment; filename="invoice-1042.pdf"'
          content:
            application/pdf:
              schema:
                type: string
                format: binary
        '404':
          $ref: '#/components/responses/NotFound'
```

`format: binary` tells generators to treat the body as raw bytes (a `Blob`, `ReadableStream`, `byte[]`, or `Buffer` depending on language) rather than parsing it. Use `format: byte` only for actual base64-encoded content embedded in JSON, which is rare for downloads; a normal file response is `binary`.

## Common download media types

| Artifact      | Media type                                                          | Notes                                      |
|---------------|---------------------------------------------------------------------|--------------------------------------------|
| PDF           | `application/pdf`                                                   | Invoices, reports, contracts               |
| CSV           | `text/csv`                                                          | Tabular exports; declare charset if needed |
| Excel         | `application/vnd.openxmlformats-officedocument.spreadsheetml.sheet` | `.xlsx`                                    |
| ZIP           | `application/zip`                                                   | Bundled assets, bulk export                |
| PNG/JPEG/WebP | `image/png`, `image/jpeg`, `image/webp`                             | Generated or resized images                |
| Plain text    | `text/plain`                                                        | Logs, `.env` templates                     |
| Unknown / any | `application/octet-stream`                                          | Fallback; prefer a specific type           |

For CSV, state the dialect in the description when it matters (comma vs semicolon, header row, UTF-8 with BOM for Excel, date formatting). Clients parsing the file cannot infer these from the media type alone.

## inline versus attachment

`Content-Disposition` decides whether the browser renders the file or downloads it:

-   `inline; filename="invoice.pdf"` lets the browser display it in a tab (common for PDF and images).
-   `attachment; filename="invoice-1042.pdf"` forces a Save dialog and suggests the file name.

Document the header as part of the response and, where the client can choose, accept a query parameter:

``` yaml
parameters:
  - name: download
    in: query
    schema: { type: boolean, default: false }
    description: When true, sets Content-Disposition to attachment to force a download.
```

Always provide a stable filename, ASCII-safe or RFC 5987 encoded for non-ASCII names, so saved files are not named after the endpoint.

## Multiple formats from one endpoint

An export that can be CSV, XLSX, or JSON is a content-negotiation decision. Use either an `Accept` header or an explicit format parameter and list each success representation:

``` yaml
responses:
  '200':
    description: The export in the requested format.
    content:
      text/csv:
        schema: { type: string, format: binary }
      application/vnd.openxmlformats-officedocument.spreadsheetml.sheet:
        schema: { type: string, format: binary }
      application/json:
        schema:
          type: array
          items: { $ref: '#/components/schemas/Order' }
  '406':
    description: Requested format is not supported.
```

Listing each media type lets a generator expose per-format methods and lets documentation render the right example. Do not collapse them into `application/octet-stream`, which hides the real contract.

## Errors are JSON, not the file

A download endpoint can still fail with a structured body. Give error responses their own JSON media type so clients do not attempt to parse an HTML proxy error page or a JSON error as the file:

``` yaml
'401':
  description: Authentication required.
  content:
    application/json:
      schema: { $ref: '#/components/schemas/ProblemDetail' }
'403':
  description: The file exists but this account may not access it.
  content:
    application/json:
      schema: { $ref: '#/components/schemas/ProblemDetail' }
```

Clients must check the response status and `Content-Type` before treating bytes as the expected file; a `200` with `application/json` after an auth redirect is a common source of a saved "PDF" that is actually a login page.

## Resumable and ranged downloads

Large files should support range requests and conditional GET. Document the headers so download managers and clients can resume and cache:

``` yaml
'200':
  headers:
    Accept-Ranges:
      schema: { type: string, example: bytes }
    ETag:
      schema: { type: string }
    Last-Modified:
      schema: { type: string, format: http-date }
  content:
    application/zip:
      schema: { type: string, format: binary }
'206':
  description: Partial content for a Range request.
  headers:
    Content-Range:
      schema: { type: string, example: bytes 0-1048575/8388608 }
  content:
    application/zip:
      schema: { type: string, format: binary }
'304':
  description: Not modified; use the cached copy.
```

`If-None-Match`/`ETag` and `If-Modified-Since` avoid re-downloading unchanged files; `Range`/`206 Partial Content` enables resume and parallel fetch. Signed URLs that expire should return a documented `403` or `410` when stale, with a documented way to fetch a fresh link.

## Async-generated files

A file that takes time to generate should not block. Kick off a job that returns `202` and a job resource; when the job succeeds, its result is a short-lived signed download URL modeled exactly as above. This composes the async-job pattern with the download pattern instead of holding the request open for minutes.

## What codegen, mocks, and AI callers need

-   `format: binary` is the signal that makes a generator return a `Blob`/`byte[]`; without it, JavaScript clients corrupt binary by decoding as text.
-   Multiple `content` entries let typed clients request a specific representation and validate `406`.
-   A spec-driven mock should return a small valid fixture (a real one-page PDF or a few CSV lines) with the correct headers, plus a JSON `404`, so the client's save-to-disk and error paths are both testable.
-   An AI agent wiring up a download needs the media type, the `Content-Disposition` behavior, and the fact that errors are JSON; without that it will guess the extension and mishandle failures.

## Checklist

1.  Model file bodies as `type: string, format: binary` with the precise media type; never as an object or plain string.
2.  Document `Content-Disposition`, including a safe filename and inline versus attachment.
3.  List each supported format as a separate `content` entry and add a `406`.
4.  Give errors a JSON Problem Detail body and tell clients to verify status and Content-Type before saving.
5.  Add `Accept-Ranges`, `ETag`, `Last-Modified`, `206`, and `304` for large or cacheable files.
6.  Document signed-URL expiry and how to obtain a fresh link.
7.  Generate the client and confirm downloads arrive as bytes, and mock both a valid file and a JSON error.

Get these right and invoices export without corruption, large bundles resume cleanly, and a failed download shows a real error instead of a PDF full of HTML.

You can describe binary responses, generate byte-accurate clients, and mock a real file plus a JSON error in one local-first workspace, [right in your browser](https://www.powerduck.com/app/?ref=powerduck.com). For the upload side of the same contract, see [file uploads and multipart/form-data in OpenAPI](https://www.powerduck.com/blog/openapi-file-upload-multipart-form-data-guide/).

