Rate limits
How many calls an agent may run and start, how to read the waits and how to stay under.
Limits are per agent (or per account), not per key: more keys don't add room. A refused call is always free and never counts against the limits.
- Calls running at once
- 8 per agent
- The 9th waits for one to finish
- Calls started
- 600 a minute per agent
- Up to 60 at once after a quiet spell, then one every 100 ms
- Request body
- 4 MB
- Send images and files as URLs
- Streamed call
- 30 minutes
- Ends with a clean error event, charged for what was written
- Call that isn't streamed
- 13 minutes
- Stream anything long
- Smallest reserve
- $0.001
- Every call holds at least this while it runs
- Active keys
- 10 per agent or account
- Revoke one to make room
- Key spending limit
- $0.01 to $1,000,000
- Per day, week or month, reset at 00:00 UTC
- Waits we name
- 1 to 60 s
- In Retry-After and Retry-After-Ms, plus up to a fifth at random
Try it
Start some calls.
The two limits you can meet
8 calls running at once. The 9th is refused with 429 too_many_running_calls and a wait of 2 to 8 seconds, spread so a refused batch doesn't return all at once. A slot frees the moment a call ends.
600 calls started a minute. Starts are spaced on a schedule: one every 100 ms, with up to 60 at once after a quiet spell. Faster than that is refused with 429 rate_limited and the exact wait.
With calls that take a second or more, the 8-call limit is the one you meet first. The rate matters for many very short calls.
Reading the headers
Every 429 and 503 names its wait:
HTTP/1.1 429 Too Many Requests
retry-after: 2
retry-after-ms: 1180retry-after-msis exact;retry-afteris the same in whole seconds.- Waits run from 1 second to 60. We add up to a fifth at random so clients told the same wait don't all come back at once.
- The OpenAI and Anthropic SDKs read these headers and wait on their own.
Busy days
Every agent shares the day's AI capacity. When 80% of it is used, an agent that already had its fair share today waits until 00:00 UTC with 429 daily_share_used, so the rest stays for everyone. It's rare, and it lifts sooner when capacity is added.
Stay under
- Queue on your side. Run at most 8 calls per agent at once. A simple semaphore does it.
- Stream. Streamed calls hold a slot as long as other calls, but fail less on slow networks.
- Let the SDK retry. It already honors
Retry-After-Ms. See Retries.