Errors

Every error the API returns, why it happens, what to do and whether to retry.

Every error carries a stable code in the body, in the format's own shape. Branch on the code, not on the message: messages may be reworded.

31 of 31
401invalid_api_key
When
The key is missing, mistyped, unknown or revoked. A key revoked while a call starts is refused at its reserve.
What to do
Send the full key as Authorization: Bearer binf_... or x-api-key: binf_.... If it was revoked, make a new one.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "Missing or malformed API key. Send your binf_ key as `Authorization: Bearer <key>` or `x-api-key`.",    "type": "authentication_error",    "code": "invalid_api_key",    "param": null  }}
Messages (HTTP 401)
{  "type": "error",  "error": {    "type": "authentication_error",    "message": "Missing or malformed API key. Send your binf_ key as `Authorization: Bearer <key>` or `x-api-key`.",    "code": "invalid_api_key"  }}
403agent_inactive
When
The agent's launch is not confirmed on chain yet, or its launch fee is unpaid.
What to do
Finish the launch on the agent's page. Keys work once it says the agent is live.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "This agent is not active yet. Activate it on its page to use the gateway.",    "type": "permission_error",    "code": "agent_inactive",    "param": null  }}
Messages (HTTP 403)
{  "type": "error",  "error": {    "type": "permission_error",    "message": "This agent is not active yet. Activate it on its page to use the gateway.",    "code": "agent_inactive"  }}
403agent_suspended
When
The agent was suspended. None of its keys can start calls.
What to do
Contact bInference.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "This agent is suspended.",    "type": "permission_error",    "code": "agent_suspended",    "param": null  }}
Messages (HTTP 403)
{  "type": "error",  "error": {    "type": "permission_error",    "message": "This agent is suspended.",    "code": "agent_suspended"  }}
400invalid_request
When
The body is not JSON, or misses a field the format needs, such as model or messages. A model's provider can refuse a parameter the same way.
What to do
Check the body against the endpoint's reference, and the parameters against the model.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "The request body is not valid JSON.",    "type": "invalid_request_error",    "code": "invalid_request",    "param": null  }}
Messages (HTTP 400)
{  "type": "error",  "error": {    "type": "invalid_request_error",    "message": "The request body is not valid JSON.",    "code": "invalid_request"  }}
400unknown_model
When
No listed model has this id.
What to do
Copy an id from GET /models or the models page. Ids look like anthropic/claude-sonnet-5.5.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "Unknown model `openai/gpt-9`. Every model we serve is listed at /api/v1/models.",    "type": "invalid_request_error",    "code": "unknown_model",    "param": null  }}
Messages (HTTP 400)
{  "type": "error",  "error": {    "type": "invalid_request_error",    "message": "Unknown model `openai/gpt-9`. Every model we serve is listed at /api/v1/models.",    "code": "unknown_model"  }}
400model_not_served
When
Free models, :free and :batch versions, retired versions and routers that pick a model after the call. Their price can't be reserved up front.
What to do
Use the paid model's own id, or add :nitro, :floor, :exacto or :online to it.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "`meta-llama/llama-5:free` is a free model, and free models are not served here. Every model we serve is listed at /api/v1/models.",    "type": "invalid_request_error",    "code": "model_not_served",    "param": null  }}
Messages (HTTP 400)
{  "type": "error",  "error": {    "type": "invalid_request_error",    "message": "`meta-llama/llama-5:free` is a free model, and free models are not served here. Every model we serve is listed at /api/v1/models.",    "code": "model_not_served"  }}
400tool_not_served
When
Tools that run other models, or that keep files, memory or a sandbox on a shared account: image generation, hosted code execution, file search, memory, X search. Also a tool loop with no step limit.
What to do
Remove the tool, or run it on your own machine. Web search and web fetch are served. Set max_tool_calls on tool loops.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "X search bills for every post it returns, with no limit a request can set, so its cost cannot be reserved up front. Remove `x_search` to search the web only.",    "type": "invalid_request_error",    "code": "tool_not_served",    "param": null  }}
Messages (HTTP 400)
{  "type": "error",  "error": {    "type": "invalid_request_error",    "message": "X search bills for every post it returns, with no limit a request can set, so its cost cannot be reserved up front. Remove `x_search` to search the web only.",    "code": "tool_not_served"  }}
413request_too_large
When
The body is over 4 MB, or the prompt is longer than the model's context.
What to do
Send images and files as URLs, not inline base64. In Claude Code, run /compact or /clear.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "The request body is larger than 4 MB. Send images and files as URLs instead of inline data.",    "type": "invalid_request_error",    "code": "request_too_large",    "param": null  }}
Messages (HTTP 413)
{  "type": "error",  "error": {    "type": "request_too_large",    "message": "The request body is larger than 4 MB. Send images and files as URLs instead of inline data.",    "code": "request_too_large"  }}
400request_over_capacity
When
One request's reserve is larger than what the gateway can hold for a single call at the moment.
What to do
Lower max_tokens (max_output_tokens on Responses), then send it again.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "This request asks for more output than the gateway can take on right now. Lower max_tokens (max_output_tokens on Responses) and retry.",    "type": "invalid_request_error",    "code": "request_over_capacity",    "param": null  }}
Messages (HTTP 400)
{  "type": "error",  "error": {    "type": "invalid_request_error",    "message": "This request asks for more output than the gateway can take on right now. Lower max_tokens (max_output_tokens on Responses) and retry.",    "code": "request_over_capacity"  }}
404not_found
When
The path doesn't exist, such as embeddings or token counting.
What to do
Use one of the five endpoints in this reference.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "`POST /v1/embeddings` is not served here. The models are served at /v1/chat/completions, /v1/responses and /v1/messages.",    "type": "not_found_error",    "code": "not_found",    "param": null  }}
Messages (HTTP 404)
{  "type": "error",  "error": {    "type": "not_found_error",    "message": "`POST /v1/embeddings` is not served here. The models are served at /v1/chat/completions, /v1/responses and /v1/messages.",    "code": "not_found"  }}
402insufficient_balance
When
What the agent can spend now is less than this call's reserve, its worst-case cost. The message names both amounts.
What to do
Lower max_tokens so the reserve fits, wait for running calls to finish or add credit. Agents earn credit from their token's trading fees.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "Balance $0.004210 ($0.001000 held by calls running or being priced) is below this request's reserve of $0.019660. Lower the output limit or add credit.",    "type": "billing_error",    "code": "insufficient_balance",    "param": null  }}
Messages (HTTP 402)
{  "type": "error",  "error": {    "type": "billing_error",    "message": "Balance $0.004210 ($0.001000 held by calls running or being priced) is below this request's reserve of $0.019660. Lower the output limit or add credit.",    "code": "insufficient_balance"  }}
402key_limit
When
The key's own daily, weekly or monthly limit can't cover this call's reserve. The agent's other keys keep working.
What to do
Wait for the reset named in the message, lower max_tokens or raise the limit on the keys page.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "This key's limit of $5.000000 a day is used up: $4.990000 spent and $0.000000 held by calls running or being priced, and this request reserves $0.019660. It resets at 2026-10-01T00:00:00.000Z. Lower the output limit, or ask the agent's owner to raise the key's limit.",    "type": "billing_error",    "code": "key_limit",    "param": null  }}
Messages (HTTP 402)
{  "type": "error",  "error": {    "type": "billing_error",    "message": "This key's limit of $5.000000 a day is used up: $4.990000 spent and $0.000000 held by calls running or being priced, and this request reserves $0.019660. It resets at 2026-10-01T00:00:00.000Z. Lower the output limit, or ask the agent's owner to raise the key's limit.",    "code": "key_limit"  }}
429rate_limited
When
The agent started more than 60 calls at once, or more than 600 a minute.
What to do
Wait for Retry-After-Ms and send again. Official OpenAI and Anthropic SDKs do this for you.
Retry
Retry after the wait in Retry-After-Ms
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "This agent is starting calls faster than 600 a minute. Retry in 1 second.",    "type": "rate_limit_error",    "code": "rate_limited",    "param": null  }}
Messages (HTTP 429)
{  "type": "error",  "error": {    "type": "rate_limit_error",    "message": "This agent is starting calls faster than 600 a minute. Retry in 1 second.",    "code": "rate_limited"  }}
429too_many_running_calls
When
The agent has 8 calls running. A slot frees the moment one ends.
What to do
Queue calls on your side, or wait the 2 to 8 seconds the headers name.
Retry
Retry after the wait in Retry-After-Ms
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "This agent already has 8 calls running. Retry when one finishes.",    "type": "rate_limit_error",    "code": "too_many_running_calls",    "param": null  }}
Messages (HTTP 429)
{  "type": "error",  "error": {    "type": "rate_limit_error",    "message": "This agent already has 8 calls running. Retry when one finishes.",    "code": "too_many_running_calls"  }}
429daily_share_used
When
Rare. On a day when 80% of the shared capacity is spent, an agent that already used its fair share waits, so the rest stays for everyone.
What to do
Wait until 00:00 UTC. It often lifts sooner, when capacity is added.
Retry
Retry after the wait in Retry-After-Ms
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "Today's shared AI capacity is running low, and this agent has used its fair share of it ($12.000000 of $12.000000). It can start calls again at 00:00 UTC, or sooner if capacity is added.",    "type": "rate_limit_error",    "code": "daily_share_used",    "param": null  }}
Messages (HTTP 429)
{  "type": "error",  "error": {    "type": "rate_limit_error",    "message": "Today's shared AI capacity is running low, and this agent has used its fair share of it ($12.000000 of $12.000000). It can start calls again at 00:00 UTC, or sooner if capacity is added.",    "code": "daily_share_used"  }}
503/529model_busy
When
Every provider of this model is at capacity. On /messages it is 529 overloaded_error, so Claude Code moves to its fallback model.
What to do
Wait for Retry-After-Ms, or use another model. Listing fallbacks in models lets a call run on the next one.
Retry
Retry after the wait in Retry-After-Ms
Cost
Free
Chat Completions, Responses
{  "error": {    "message": "The model `openai/gpt-6.1` is busy at its providers right now. Retry in 12 seconds.",    "type": "server_error",    "code": "model_busy",    "param": null  }}
Messages (HTTP 529)
{  "type": "error",  "error": {    "type": "overloaded_error",    "message": "The model `openai/gpt-6.1` is busy at its providers right now. Retry in 12 seconds.",    "code": "model_busy"  }}
503gateway_busy
When
The gateway is at its capacity for a moment.
What to do
Wait for Retry-After-Ms and send again.
Retry
Retry after the wait in Retry-After-Ms
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "The gateway is busy right now. Retry in 8 seconds.",    "type": "server_error",    "code": "gateway_busy",    "param": null  }}
Messages (HTTP 503)
{  "type": "error",  "error": {    "type": "overloaded_error",    "message": "The gateway is busy right now. Retry in 8 seconds.",    "code": "gateway_busy"  }}
503gateway_capacity
When
The gateway is topping up its capacity. bInference is alerted at once.
What to do
Wait for Retry-After-Ms and send again.
Retry
Retry after the wait in Retry-After-Ms
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "The gateway is briefly out of capacity. Try again soon.",    "type": "server_error",    "code": "gateway_capacity",    "param": null  }}
Messages (HTTP 503)
{  "type": "error",  "error": {    "type": "overloaded_error",    "message": "The gateway is briefly out of capacity. Try again soon.",    "code": "gateway_capacity"  }}
503gateway_unavailable
When
The gateway can't reach its models for a moment. bInference is alerted at once.
What to do
Wait for Retry-After-Ms and send again.
Retry
Retry after the wait in Retry-After-Ms
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "The gateway is briefly unavailable. Try again soon.",    "type": "server_error",    "code": "gateway_unavailable",    "param": null  }}
Messages (HTTP 503)
{  "type": "error",  "error": {    "type": "overloaded_error",    "message": "The gateway is briefly unavailable. Try again soon.",    "code": "gateway_unavailable"  }}
503catalog_unavailable
When
The gateway couldn't read model prices, so it can't size the reserve.
What to do
Send again after 5 seconds.
Retry
Retry after the wait in Retry-After-Ms
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "Model prices are unavailable. Try again shortly.",    "type": "server_error",    "code": "catalog_unavailable",    "param": null  }}
Messages (HTTP 503)
{  "type": "error",  "error": {    "type": "overloaded_error",    "message": "Model prices are unavailable. Try again shortly.",    "code": "catalog_unavailable"  }}
503service_unavailable
When
The gateway couldn't reach its database before the reserve. Nothing was reserved.
What to do
Send again after about 2 seconds.
Retry
Retry after the wait in Retry-After-Ms
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "The balance is briefly unavailable. Try again in a moment.",    "type": "server_error",    "code": "service_unavailable",    "param": null  }}
Messages (HTTP 503)
{  "type": "error",  "error": {    "type": "overloaded_error",    "message": "The balance is briefly unavailable. Try again in a moment.",    "code": "service_unavailable"  }}
500gateway_error
When
A fault in the gateway after the reserve. The reserve is released and the call is free.
What to do
Send it again. If it keeps failing, contact bInference with the x-binference-call-id header.
Retry
Retry
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "The gateway failed on this call. Retry it.",    "type": "server_error",    "code": "gateway_error",    "param": null  }}
Messages (HTTP 500)
{  "type": "error",  "error": {    "type": "api_error",    "message": "The gateway failed on this call. Retry it.",    "code": "gateway_error"  }}
504upstream_timeout
When
The model sent nothing before the time limit: 30 minutes streamed, 13 minutes not streamed.
What to do
Stream long calls, lower max_tokens or use a faster model.
Retry
Retry
Cost
Only what the provider recorded, usually nothing
Chat Completions, Responses, models, balance
{  "error": {    "message": "The model did not answer in time.",    "type": "server_error",    "code": "upstream_timeout",    "param": null  }}
Messages (HTTP 504)
{  "type": "error",  "error": {    "type": "api_error",    "message": "The model did not answer in time.",    "code": "upstream_timeout"  }}
502upstream_unreachable
When
The connection to the model's provider failed.
What to do
Send it again.
Retry
Retry
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "Could not reach the model provider. Try again.",    "type": "server_error",    "code": "upstream_unreachable",    "param": null  }}
Messages (HTTP 502)
{  "type": "error",  "error": {    "type": "api_error",    "message": "Could not reach the model provider. Try again.",    "code": "upstream_unreachable"  }}
502provider_error
When
The model's provider failed, or refused this one call.
What to do
Send it again, or use another model.
Retry
Retry
Cost
Only what the provider recorded, usually nothing
Chat Completions, Responses, models, balance
{  "error": {    "message": "The model's provider failed to answer. Try again in a moment.",    "type": "server_error",    "code": "provider_error",    "param": null  }}
Messages (HTTP 502)
{  "type": "error",  "error": {    "type": "api_error",    "message": "The model's provider failed to answer. Try again in a moment.",    "code": "provider_error"  }}
503provider_unavailable
When
The model is down at its providers.
What to do
Send it again shortly, or use another model.
Retry
Retry
Cost
Only what the provider recorded, usually nothing
Chat Completions, Responses, models, balance
{  "error": {    "message": "The model is unavailable right now. Try again in a moment.",    "type": "server_error",    "code": "provider_unavailable",    "param": null  }}
Messages (HTTP 503)
{  "type": "error",  "error": {    "type": "overloaded_error",    "message": "The model is unavailable right now. Try again in a moment.",    "code": "provider_unavailable"  }}
404no_provider
When
No provider of the model supports every option in the request, such as tools, images or a long context.
What to do
Drop an option, or pick a model that lists it.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "No provider can serve this request with this model. Try another model, or fewer options such as tools or images.",    "type": "not_found_error",    "code": "no_provider",    "param": null  }}
Messages (HTTP 404)
{  "type": "error",  "error": {    "type": "not_found_error",    "message": "No provider can serve this request with this model. Try another model, or fewer options such as tools or images.",    "code": "no_provider"  }}
403provider_refused
When
The provider's own policy refused the request.
What to do
Change the request, or use another model.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "The model's provider refused this request for this model.",    "type": "permission_error",    "code": "provider_refused",    "param": null  }}
Messages (HTTP 403)
{  "type": "error",  "error": {    "type": "permission_error",    "message": "The model's provider refused this request for this model.",    "code": "provider_refused"  }}
422unprocessable
When
The provider accepted the request but couldn't run it.
What to do
Check the parameters against the model, or use another model.
Retry
Fix first
Cost
Free
Chat Completions, Responses, models, balance
{  "error": {    "message": "The model's provider could not process this request.",    "type": "server_error",    "code": "unprocessable",    "param": null  }}
Messages (HTTP 422)
{  "type": "error",  "error": {    "type": "api_error",    "message": "The model's provider could not process this request.",    "code": "unprocessable"  }}
408timeout
When
The provider gave up on the call (408 or 504).
What to do
Send it again, stream it or lower max_tokens.
Retry
Retry
Cost
Only what the provider recorded, usually nothing
Chat Completions, Responses, models, balance
{  "error": {    "message": "The model took too long to answer. Try again.",    "type": "server_error",    "code": "timeout",    "param": null  }}
Messages (HTTP 408)
{  "type": "error",  "error": {    "type": "api_error",    "message": "The model took too long to answer. Try again.",    "code": "timeout"  }}
502upstream_error
When
Any other error from the provider.
What to do
Send it again shortly.
Retry
Retry
Cost
Only what the provider recorded, usually nothing
Chat Completions, Responses, models, balance
{  "error": {    "message": "The request could not be completed. Try again in a moment.",    "type": "server_error",    "code": "upstream_error",    "param": null  }}
Messages (HTTP 502)
{  "type": "error",  "error": {    "type": "api_error",    "message": "The request could not be completed. Try again in a moment.",    "code": "upstream_error"  }}

The short version

StatusMeansRetry?
400The request can't be served as sentNo. Fix it first
401The key is missing, wrong or revokedNo
402The balance or the key's limit can't cover the reserveNo. Lower max_tokens or add credit
403The agent isn't activeNo
404Unknown path, or no provider for this requestNo
413The body is over 4 MBNo. Send files as URLs
429Too many running calls, or too fastYes, after Retry-After-Ms
500The gateway failed. The call is freeYes
502 504The model's provider failedYes
503 529Busy for a momentYes, after Retry-After-Ms

Errors inside a stream

Once an answer has started streaming, an error can't change the status any more. It arrives as the format's own error event, so SDKs raise it instead of hanging:

data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"error"}],"error":{"message":"The model took too long to answer. Try again.","type":"server_error","code":"timeout"}}

The call is charged for what was written before the error.

On this page