Rate limits

How many calls an agent may run and start, how to read the waits and how to stay under.

Limits are per agent (or per account), not per key: more keys don't add room. A refused call is always free and never counts against the limits.

Calls running at once
8 per agent
The 9th waits for one to finish
Calls started
600 a minute per agent
Up to 60 at once after a quiet spell, then one every 100 ms
Request body
4 MB
Send images and files as URLs
Streamed call
30 minutes
Ends with a clean error event, charged for what was written
Call that isn't streamed
13 minutes
Stream anything long
Smallest reserve
$0.001
Every call holds at least this while it runs
Active keys
10 per agent or account
Revoke one to make room
Key spending limit
$0.01 to $1,000,000
Per day, week or month, reset at 00:00 UTC
Waits we name
1 to 60 s
In Retry-After and Retry-After-Ms, plus up to a fifth at random

Try it

Try the limits
Each call lasts
Calls running0 of 8
Can start at once60 of 60
Started
0
Hit 8 running
0
Hit the rate
0

Start some calls.

This runs the gateway's own rules in your browser. An agent may have 8 calls running and start 600 a minute: up to 60 at once when rested, then one every 100 ms. With calls longer than about a second, the 8-call cap is the one you meet. A refused call costs nothing and doesn't use up the allowance.

The two limits you can meet

8 calls running at once. The 9th is refused with 429 too_many_running_calls and a wait of 2 to 8 seconds, spread so a refused batch doesn't return all at once. A slot frees the moment a call ends.

600 calls started a minute. Starts are spaced on a schedule: one every 100 ms, with up to 60 at once after a quiet spell. Faster than that is refused with 429 rate_limited and the exact wait.

With calls that take a second or more, the 8-call limit is the one you meet first. The rate matters for many very short calls.

Reading the headers

Every 429 and 503 names its wait:

HTTP/1.1 429 Too Many Requests
retry-after: 2
retry-after-ms: 1180
  • retry-after-ms is exact; retry-after is the same in whole seconds.
  • Waits run from 1 second to 60. We add up to a fifth at random so clients told the same wait don't all come back at once.
  • The OpenAI and Anthropic SDKs read these headers and wait on their own.

Busy days

Every agent shares the day's AI capacity. When 80% of it is used, an agent that already had its fair share today waits until 00:00 UTC with 429 daily_share_used, so the rest stays for everyone. It's rare, and it lifts sooner when capacity is added.

Stay under

  • Queue on your side. Run at most 8 calls per agent at once. A simple semaphore does it.
  • Stream. Streamed calls hold a slot as long as other calls, but fail less on slow networks.
  • Let the SDK retry. It already honors Retry-After-Ms. See Retries.

On this page