---
title: Rate limits
description: Per-key RPM and concurrency limits, the RateLimit-* response headers, and how to handle a 429 response.
---

Each API key is capped at **120 requests per minute (RPM)** and **10 concurrent generations** by default; both values are configurable. When you exceed the RPM cap, the request is rejected with `429 rate_limit_exceeded` and a `Retry-After` header telling you how many seconds to wait before retrying.

## Default limits

<TypeTable
  type={{
    'RPM (requests/min)': {
      type: 'number',
      description: 'Defaults to 120 per key. Configurable. Exceeding it returns 429 rate_limit_exceeded.',
    },
    'Concurrent generations': {
      type: 'number',
      description: 'Defaults to 10 in-flight generations per key. Configurable, independent of the RPM cap.',
    },
  }}
/>

## Response headers

The `RateLimit-*` headers are returned on every response, so you can track your remaining budget without waiting for a `429`.

<TypeTable
  type={{
    'RateLimit-Limit': {
      type: 'integer',
      description: 'The current RPM limit for the key.',
    },
    'RateLimit-Remaining': {
      type: 'integer',
      description: 'How many requests remain in the current window.',
    },
    'RateLimit-Reset': {
      type: 'integer (seconds)',
      description: 'Seconds until the window resets and the counter clears.',
    },
    'Retry-After': {
      type: 'integer (seconds)',
      description: 'Sent only on a 429: how many seconds to wait before retrying.',
    },
  }}
/>

## The 429 response

```http
HTTP/1.1 429 Too Many Requests
RateLimit-Limit: 120
RateLimit-Remaining: 0
RateLimit-Reset: 37
Retry-After: 37
```

```json
{
  "error": {
    "type": "rate_limit_error",
    "code": "rate_limit_exceeded",
    "message": "Rate limit exceeded. Please retry later."
  }
}
```

<Callout type="info" title="Honour Retry-After">
On a `429`, wait exactly the number of seconds given in `Retry-After` (or `RateLimit-Reset`) before retrying. That is more reliable than a fixed delay.
</Callout>

<Callout type="warn" title="Concurrency is capped separately">
The 10 concurrent-generation limit is independent of the RPM cap. If you run many long video jobs, queue them on your side so you don't hit the concurrency ceiling.
</Callout>
