# Designing a search endpoint in OpenAPI: q, filters, sorting, sparse fieldsets, and pagination codegen understands

Every list endpoint grows search until it becomes a dialect of its own. One team uses `status_in=paid,refunded`, another uses `filter[status]=paid`, a third invents `statuses=...`; sorting is a `sort` string whose allowed values you discover by reading the controller; full text is sometimes `search`, sometimes `q`, sometimes `name`. Clients cannot generate a typed SDK against a dialect. A small set of consistent parameters, each with a documented type, turns search into something codegen can model.

## Reserve a small vocabulary

| Parameter               | Job                                                        | Example                       |
|-------------------------|------------------------------------------------------------|-------------------------------|
| `q`                     | Free-text search across the resource's searchable fields   | `q=ceramic knife`             |
| `filter` / typed params | Structured constraints on specific fields                  | `status=paid&amount_min=1000` |
| `sort`                  | Ordered, comma-separated fields; `-` prefix for descending | `sort=-created_at,name`       |
| `fields`                | Sparse fieldset, limit returned fields                     | `fields=id,name,amount`       |
| `page` / cursor         | Pagination (covered separately)                            | `page[cursor]=eyc...`         |

Keep `q` for full text only; do not overload it with structured predicates. If a constraint has a type and a finite set of operators, model it as a filter, not as part of the query string.

## Filters: prefer explicit typed parameters

The most codegen-friendly style is one parameter per constraint with a real schema, because generators then emit typed arguments and validators know the type:

``` yaml
parameters:
  - name: q
    in: query
    schema: { type: string }
    description: Free-text match over name and email.
  - name: status
    in: query
    schema:
      type: array
      items:
        type: string
        enum: [paid, refunded, failed, pending]
    style: form
    explode: false
    description: Match any of the given statuses; repeat or comma-separate.
  - name: amount_min
    in: query
    schema: { type: integer }
    description: Lower bound on amount in minor units, inclusive.
  - name: amount_max
    in: query
    schema: { type: integer }
  - name: created_after
    in: query
    schema: { type: string, format: date-time }
```

`explode: false` renders the array as `status=paid,refunded`; the default `explode: true` renders `status=paid&status=refunded`. Pick one and document it; clients and generated SDKs differ in how they encode arrays, and an undocumented choice is a recurring bug.

For a truly dynamic filter surface (a query builder over dozens of fields), the bracket convention `filter[field]=value` keeps everything under one namespace and avoids colliding with reserved params, but it costs you static typing; generators treat it as an open map. Use typed parameters for the stable, indexed fields your UI actually exposes and reserve the dynamic form for ad-hoc admin tooling.

## Operators need a documented grammar

Ranges, negation, and partial match should not be smuggled into magic strings. Choose explicit parameter names (`amount_min`/`amount_max`, `created_after`/`created_before`) for the common cases. If you must support many operators per field, define a small grammar and publish it rather than letting each client invent one:

``` text
filter=<field>:<operator>:<value>
operators: eq, neq, gt, gte, lt, lte, in, nin, like, isnull
example: filter=amount:gte:1000
```

Whatever the grammar, keep it identical across resources, reject unknown fields and operators with a clear `400` listing the allowed set, and never silently ignore an unsupported filter (silently returning unfiltered data is worse than an error).

## Sorting with a sign convention

Use a single `sort` parameter with comma-separated fields and a leading `-` for descending, matching the JSON:API convention:

``` yaml
- name: sort
  in: query
  schema:
    type: string
    pattern: '^-?[a-z_]+(,-?[a-z_]+)*$'
    example: -created_at,name
  description: >-
    Comma-separated sortable fields; prefix a field with '-' for descending.
    Allowed fields: created_at, amount, name, status. Unknown fields return 400.
```

Whitelist the sortable fields server-side; an arbitrary sort field is both a correctness and a SQL-injection surface. Document the tiebreaker (for example a unique `id` as the final sort key) so pagination is stable when two rows share the primary sort value.

## Sparse fieldsets for payload size

List endpoints often return huge objects to every caller. A `fields` parameter lets clients request only what they render:

``` yaml
- name: fields
  in: query
  schema:
    type: array
    items: { type: string }
    explode: false
  description: Comma-separated field names to return; omitted returns the default set.
  example: id,name,amount,status
```

Treat an omitted `fields` as a documented default set (not necessarily every field), validate requested names against the schema, and reject unknown names rather than ignoring them. Sparse fieldsets also let you keep expensive nested relations out of the default list response while still exposing them on demand.

## The response envelope

Return the same collection envelope your other list endpoints use, with the echoed query, a total where it is cheap, and pagination. Do not invent a search-specific response shape:

``` yaml
SearchCustomersResponse:
  type: object
  required: [data]
  properties:
    data:
      type: array
      items:
        $ref: '#/components/schemas/CustomerSummary'
    page:
      $ref: '#/components/schemas/PageInfo'
    query:
      type: object
      description: Echoed, normalized search parameters for debugging and sharing.
      properties:
        q: { type: string }
        status:
          type: array
          items: { type: string }
        sort: { type: string }
```

Echoing the normalized query makes a search result link shareable and shows the client exactly how ambiguous input was interpreted. For large or unbounded totals, prefer cursor pagination and omit an exact `total` rather than returning a number that is expensive or wrong; document that choice.

## Consistency rules that pay off

-   Same parameter names and operators on every resource; a client that learns search once should not relearn it per endpoint.
-   Types come from real schemas (dates are `date-time`, amounts are integers, enums are enums), so generated SDKs accept typed values instead of opaque strings.
-   Unknown filters, sort fields, and fieldset names are a `400` with the allowed list, never a silent no-op.
-   Full text (`q`) is fuzzy and ranked; structured filters are exact. Keep them separate so callers know what to expect.
-   Document which fields are searchable, sortable, and filterable; a parameter that exists but is not indexed is a performance incident waiting to happen.

## What codegen, mocks, and AI callers need

Typed query parameters let a generator produce method signatures like `searchCustomers({ q, status, amountMin, sort, fields })` with enums and number types; a bracket-style dynamic filter collapses to `Record<string, string>` and loses all of that. A spec-driven mock can return a filtered, sorted subset deterministically when the parameters are typed, which makes the client's table filters testable. An AI agent constructing a search call will guess `search=` and invent operators unless the vocabulary (`q`, `sort`, `*_min`/`*_max`) is explicit in the spec.

## Checklist

1.  Use `q` for free text and typed parameters for structured filters; keep them separate.
2.  Give every filter a real schema type and document array encoding (`explode`) once.
3.  Define operators with explicit bound parameters or one published grammar; reject unknowns with `400`.
4.  Use one signed `sort` string, whitelist sortable fields, and document a stable tiebreaker.
5.  Add sparse `fields` with a documented default set and validation.
6.  Reuse the standard collection envelope, echo the normalized query, and document pagination.
7.  Keep names, operators, and conventions identical across resources.
8.  Generate the SDK and confirm search arguments are typed, and test filtered/sorted output against a mock.

Get these right and every list endpoint speaks the same language, the generated SDK is typed end to end, and clients stop reverse-engineering your query dialect.

You can define the search parameters, generate a typed SDK, and exercise filtered and sorted results against a mock in one local-first workspace, [right in your browser](https://www.powerduck.com/app/?ref=powerduck.com). For the cursor and page mechanics that pair with these filters, see [cursor vs offset pagination in API design](https://www.powerduck.com/blog/cursor-vs-offset-pagination-api-design/).

