Skip to main content

Command Palette

Search for a command to run...

File downloads in OpenAPI: CSV, PDF, ZIP, images, Content-Disposition, and non-JSON responses

APIs return more than JSON: CSV exports, PDF invoices, ZIP archives, resized images, and generated audio. Model them with the right media type and string/binary schema, document Content-Disposition and Range, and distinguish inline from attachment. Here is how to describe downloa

Updated
•6 min read•View as Markdown

Most OpenAPI tutorials stop at application/json, yet a large share of real endpoints hand back a file: an invoice PDF, a CSV export, a ZIP of assets, a resized image, a generated spreadsheet. Teams either leave these endpoints undocumented or describe them as type: string with no format, so a generated client tries to decode a PDF as UTF-8 text and corrupts it. Non-JSON responses need explicit media types, a binary schema, and the headers that tell the client what to do with the bytes.

The schema for a binary body

A file response uses type: string with format: binary (OpenAPI 3.x), not an object and not an untyped string. The media type says what the bytes are:

paths:
  /invoices/{invoiceId}/pdf:
    get:
      summary: Download an invoice as PDF
      operationId: downloadInvoicePdf
      parameters:
        - name: invoiceId
          in: path
          required: true
          schema: { type: string }
      responses:
        '200':
          description: The invoice PDF.
          headers:
            Content-Disposition:
              schema: { type: string }
              description: 'attachment; filename="invoice-1042.pdf"'
          content:
            application/pdf:
              schema:
                type: string
                format: binary
        '404':
          $ref: '#/components/responses/NotFound'

format: binary tells generators to treat the body as raw bytes (a Blob, ReadableStream, byte[], or Buffer depending on language) rather than parsing it. Use format: byte only for actual base64-encoded content embedded in JSON, which is rare for downloads; a normal file response is binary.

Common download media types

Artifact Media type Notes
PDF application/pdf Invoices, reports, contracts
CSV text/csv Tabular exports; declare charset if needed
Excel application/vnd.openxmlformats-officedocument.spreadsheetml.sheet .xlsx
ZIP application/zip Bundled assets, bulk export
PNG/JPEG/WebP image/png, image/jpeg, image/webp Generated or resized images
Plain text text/plain Logs, .env templates
Unknown / any application/octet-stream Fallback; prefer a specific type

For CSV, state the dialect in the description when it matters (comma vs semicolon, header row, UTF-8 with BOM for Excel, date formatting). Clients parsing the file cannot infer these from the media type alone.

inline versus attachment

Content-Disposition decides whether the browser renders the file or downloads it:

  • inline; filename="invoice.pdf" lets the browser display it in a tab (common for PDF and images).
  • attachment; filename="invoice-1042.pdf" forces a Save dialog and suggests the file name.

Document the header as part of the response and, where the client can choose, accept a query parameter:

parameters:
  - name: download
    in: query
    schema: { type: boolean, default: false }
    description: When true, sets Content-Disposition to attachment to force a download.

Always provide a stable filename, ASCII-safe or RFC 5987 encoded for non-ASCII names, so saved files are not named after the endpoint.

Multiple formats from one endpoint

An export that can be CSV, XLSX, or JSON is a content-negotiation decision. Use either an Accept header or an explicit format parameter and list each success representation:

responses:
  '200':
    description: The export in the requested format.
    content:
      text/csv:
        schema: { type: string, format: binary }
      application/vnd.openxmlformats-officedocument.spreadsheetml.sheet:
        schema: { type: string, format: binary }
      application/json:
        schema:
          type: array
          items: { $ref: '#/components/schemas/Order' }
  '406':
    description: Requested format is not supported.

Listing each media type lets a generator expose per-format methods and lets documentation render the right example. Do not collapse them into application/octet-stream, which hides the real contract.

Errors are JSON, not the file

A download endpoint can still fail with a structured body. Give error responses their own JSON media type so clients do not attempt to parse an HTML proxy error page or a JSON error as the file:

'401':
  description: Authentication required.
  content:
    application/json:
      schema: { $ref: '#/components/schemas/ProblemDetail' }
'403':
  description: The file exists but this account may not access it.
  content:
    application/json:
      schema: { $ref: '#/components/schemas/ProblemDetail' }

Clients must check the response status and Content-Type before treating bytes as the expected file; a 200 with application/json after an auth redirect is a common source of a saved "PDF" that is actually a login page.

Resumable and ranged downloads

Large files should support range requests and conditional GET. Document the headers so download managers and clients can resume and cache:

'200':
  headers:
    Accept-Ranges:
      schema: { type: string, example: bytes }
    ETag:
      schema: { type: string }
    Last-Modified:
      schema: { type: string, format: http-date }
  content:
    application/zip:
      schema: { type: string, format: binary }
'206':
  description: Partial content for a Range request.
  headers:
    Content-Range:
      schema: { type: string, example: bytes 0-1048575/8388608 }
  content:
    application/zip:
      schema: { type: string, format: binary }
'304':
  description: Not modified; use the cached copy.

If-None-Match/ETag and If-Modified-Since avoid re-downloading unchanged files; Range/206 Partial Content enables resume and parallel fetch. Signed URLs that expire should return a documented 403 or 410 when stale, with a documented way to fetch a fresh link.

Async-generated files

A file that takes time to generate should not block. Kick off a job that returns 202 and a job resource; when the job succeeds, its result is a short-lived signed download URL modeled exactly as above. This composes the async-job pattern with the download pattern instead of holding the request open for minutes.

What codegen, mocks, and AI callers need

  • format: binary is the signal that makes a generator return a Blob/byte[]; without it, JavaScript clients corrupt binary by decoding as text.
  • Multiple content entries let typed clients request a specific representation and validate 406.
  • A spec-driven mock should return a small valid fixture (a real one-page PDF or a few CSV lines) with the correct headers, plus a JSON 404, so the client's save-to-disk and error paths are both testable.
  • An AI agent wiring up a download needs the media type, the Content-Disposition behavior, and the fact that errors are JSON; without that it will guess the extension and mishandle failures.

Checklist

  1. Model file bodies as type: string, format: binary with the precise media type; never as an object or plain string.
  2. Document Content-Disposition, including a safe filename and inline versus attachment.
  3. List each supported format as a separate content entry and add a 406.
  4. Give errors a JSON Problem Detail body and tell clients to verify status and Content-Type before saving.
  5. Add Accept-Ranges, ETag, Last-Modified, 206, and 304 for large or cacheable files.
  6. Document signed-URL expiry and how to obtain a fresh link.
  7. Generate the client and confirm downloads arrive as bytes, and mock both a valid file and a JSON error.

Get these right and invoices export without corruption, large bundles resume cleanly, and a failed download shows a real error instead of a PDF full of HTML.

You can describe binary responses, generate byte-accurate clients, and mock a real file plus a JSON error in one local-first workspace, right in your browser. For the upload side of the same contract, see file uploads and multipart/form-data in OpenAPI.

More from this blog

P

Powerduck Blogs

117 posts

Essays on local-first API tooling, OpenAPI contracts, MCP, and agentic coding failures. We build Powerduck, a local-first OpenAPI studio where the spec stays the source of truth. Topics: AI code review trust, testing AI-generated code, context engineering, and what actually breaks when AI writes your code.