> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usecroma.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

> How Croma rate limits work: one allowance per organization, hourly ceilings on a few endpoints, the headers that report them, retries and 429 responses.

Most of what an organization can send is its [credit balance](/credits). This
page covers the limits on top of it.

## One allowance per organization

Limits are enforced **per organization**, not per key. Every key issued to the
same org shares one allowance, so adding keys does not multiply it.

## Hourly ceilings

A few endpoints carry an hourly ceiling on top of the plan, and job polling
draws from a bucket of its own:

| Bucket | Limit | Endpoints |
| - | - | - |
| Extract & Generate | 60 / hour on the free plan, none on paid plans | [Extract](/guides/global/extract), [Generate](/guides/global/generate). |
| Web Search | 100 / hour on the free plan, none on paid plans | [Web Search](/guides/global/web-search). |
| Research | 10 / hour on the free plan, none on paid plans | [Research](/guides/global/research). |
| Job polling | 600 / minute | [`GET /jobs/:id`](/async-jobs). Its own bucket, so polling never spends your allowance. |

On the free plan the ceilings are additional: a Research call spends one of its
10 hourly slots and 10 credits from your plan. On paid plans the credits are
the only limit. Each endpoint's page notes its limit.

## Headers

Limit state comes back as HTTP headers on every response, not in the body:

| Header | Meaning |
| - | - |
| `X-RateLimit-Limit` | The window's size: your plan's credits on data endpoints, the hourly ceiling where one applies. |
| `X-RateLimit-Remaining` | Credits, or requests in an hourly window, left before you are throttled. |
| `X-RateLimit-Reset` | ISO timestamp when the window resets. |
| `RateLimit-Policy` | The window's policy as an IETF field, for example `"credits";q=5000;w=2592000`. Present on every response, including `401` and `429`. |
| `X-Request-Id` | Unique id for the request (`req_…`). Include it in support reports. |
| `X-Cache` | `HIT` or `MISS` on cacheable endpoints. Cached hits spend no credits but still count toward an hourly ceiling. |

## Retries

Every Croma operation is a lookup: repeating a request with the same body
returns the same result and never creates or changes a record, so retrying
after a timeout or a dropped connection is always safe. An optional
`Idempotency-Key` header (any string up to 255 characters, a UUID works) comes
back in the response, so you can tie a retry to its first attempt in your logs.
Each attempt that reaches the API counts; a `429` tells you to wait for
`Retry-After` seconds rather than retry immediately.

## When you exceed a ceiling

Requests over an hourly ceiling return `429` with a `rate_limit_error`
envelope and a `Retry-After` header (seconds):

```json theme={"dark"}
{
  "error": {
    "type": "rate_limit_error",
    "code": "rate_limited",
    "message": "Rate limit exceeded. Try again in 42 seconds."
  }
}
```

```http theme={"dark"}
Retry-After: 42
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 2026-10-01T18:00:00.000Z
```

Back off until `Retry-After` elapses (or `X-RateLimit-Reset`), then retry.

<Note>
  The limiter **fails open**: if the rate-limit backend is briefly unavailable,
  requests are allowed through and no `X-RateLimit-*` headers are emitted.
  Don't depend on the headers always being present.
</Note>

<Card title="Next: Errors" icon="triangle-exclamation" href="/errors">
  The error envelope and every error code.
</Card>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.