Errors
Every error the API returns, why it happens, what to do and whether to retry.
Every error carries a stable code in the body, in the format's own shape. Branch on the code, not on the message: messages may be reworded.
31 of 31
401invalid_api_keyBad or missing keyFix first
- When
- The key is missing, mistyped, unknown or revoked. A key revoked while a call starts is refused at its reserve.
- What to do
- Send the full key as
Authorization: Bearer binf_...orx-api-key: binf_.... If it was revoked, make a new one. - Retry
- Fix first
- Cost
- Free
{ "error": { "message": "Missing or malformed API key. Send your binf_ key as `Authorization: Bearer <key>` or `x-api-key`.", "type": "authentication_error", "code": "invalid_api_key", "param": null }}{ "type": "error", "error": { "type": "authentication_error", "message": "Missing or malformed API key. Send your binf_ key as `Authorization: Bearer <key>` or `x-api-key`.", "code": "invalid_api_key" }}403agent_inactiveAgent not active yetFix first
- When
- The agent's launch is not confirmed on chain yet, or its launch fee is unpaid.
- What to do
- Finish the launch on the agent's page. Keys work once it says the agent is live.
- Retry
- Fix first
- Cost
- Free
{ "error": { "message": "This agent is not active yet. Activate it on its page to use the gateway.", "type": "permission_error", "code": "agent_inactive", "param": null }}{ "type": "error", "error": { "type": "permission_error", "message": "This agent is not active yet. Activate it on its page to use the gateway.", "code": "agent_inactive" }}403agent_suspendedAgent suspendedFix first
- When
- The agent was suspended. None of its keys can start calls.
- What to do
- Contact bInference.
- Retry
- Fix first
- Cost
- Free
{ "error": { "message": "This agent is suspended.", "type": "permission_error", "code": "agent_suspended", "param": null }}{ "type": "error", "error": { "type": "permission_error", "message": "This agent is suspended.", "code": "agent_suspended" }}400invalid_requestBad request bodyFix first
- When
- The body is not JSON, or misses a field the format needs, such as
modelormessages. A model's provider can refuse a parameter the same way. - What to do
- Check the body against the endpoint's reference, and the parameters against the model.
- Retry
- Fix first
- Cost
- Free
{ "error": { "message": "The request body is not valid JSON.", "type": "invalid_request_error", "code": "invalid_request", "param": null }}{ "type": "error", "error": { "type": "invalid_request_error", "message": "The request body is not valid JSON.", "code": "invalid_request" }}400unknown_modelUnknown modelFix first
- When
- No listed model has this id.
- What to do
- Copy an id from GET /models or the models page. Ids look like
anthropic/claude-sonnet-5.5. - Retry
- Fix first
- Cost
- Free
{ "error": { "message": "Unknown model `openai/gpt-9`. Every model we serve is listed at /api/v1/models.", "type": "invalid_request_error", "code": "unknown_model", "param": null }}{ "type": "error", "error": { "type": "invalid_request_error", "message": "Unknown model `openai/gpt-9`. Every model we serve is listed at /api/v1/models.", "code": "unknown_model" }}400model_not_servedModel not servedFix first
- When
- Free models,
:freeand:batchversions, retired versions and routers that pick a model after the call. Their price can't be reserved up front. - What to do
- Use the paid model's own id, or add
:nitro,:floor,:exactoor:onlineto it. - Retry
- Fix first
- Cost
- Free
{ "error": { "message": "`meta-llama/llama-5:free` is a free model, and free models are not served here. Every model we serve is listed at /api/v1/models.", "type": "invalid_request_error", "code": "model_not_served", "param": null }}{ "type": "error", "error": { "type": "invalid_request_error", "message": "`meta-llama/llama-5:free` is a free model, and free models are not served here. Every model we serve is listed at /api/v1/models.", "code": "model_not_served" }}400tool_not_servedTool not servedFix first
- When
- Tools that run other models, or that keep files, memory or a sandbox on a shared account: image generation, hosted code execution, file search, memory, X search. Also a tool loop with no step limit.
- What to do
- Remove the tool, or run it on your own machine. Web search and web fetch are served. Set
max_tool_callson tool loops. - Retry
- Fix first
- Cost
- Free
{ "error": { "message": "X search bills for every post it returns, with no limit a request can set, so its cost cannot be reserved up front. Remove `x_search` to search the web only.", "type": "invalid_request_error", "code": "tool_not_served", "param": null }}{ "type": "error", "error": { "type": "invalid_request_error", "message": "X search bills for every post it returns, with no limit a request can set, so its cost cannot be reserved up front. Remove `x_search` to search the web only.", "code": "tool_not_served" }}413request_too_largeRequest too largeFix first
- When
- The body is over 4 MB, or the prompt is longer than the model's context.
- What to do
- Send images and files as URLs, not inline base64. In Claude Code, run /compact or /clear.
- Retry
- Fix first
- Cost
- Free
{ "error": { "message": "The request body is larger than 4 MB. Send images and files as URLs instead of inline data.", "type": "invalid_request_error", "code": "request_too_large", "param": null }}{ "type": "error", "error": { "type": "request_too_large", "message": "The request body is larger than 4 MB. Send images and files as URLs instead of inline data.", "code": "request_too_large" }}400request_over_capacityOutput limit too high right nowFix first
- When
- One request's reserve is larger than what the gateway can hold for a single call at the moment.
- What to do
- Lower
max_tokens(max_output_tokenson Responses), then send it again. - Retry
- Fix first
- Cost
- Free
{ "error": { "message": "This request asks for more output than the gateway can take on right now. Lower max_tokens (max_output_tokens on Responses) and retry.", "type": "invalid_request_error", "code": "request_over_capacity", "param": null }}{ "type": "error", "error": { "type": "invalid_request_error", "message": "This request asks for more output than the gateway can take on right now. Lower max_tokens (max_output_tokens on Responses) and retry.", "code": "request_over_capacity" }}404not_foundPath not servedFix first
- When
- The path doesn't exist, such as embeddings or token counting.
- What to do
- Use one of the five endpoints in this reference.
- Retry
- Fix first
- Cost
- Free
{ "error": { "message": "`POST /v1/embeddings` is not served here. The models are served at /v1/chat/completions, /v1/responses and /v1/messages.", "type": "not_found_error", "code": "not_found", "param": null }}{ "type": "error", "error": { "type": "not_found_error", "message": "`POST /v1/embeddings` is not served here. The models are served at /v1/chat/completions, /v1/responses and /v1/messages.", "code": "not_found" }}402insufficient_balanceBalance too lowFix first
- When
- What the agent can spend now is less than this call's reserve, its worst-case cost. The message names both amounts.
- What to do
- Lower
max_tokensso the reserve fits, wait for running calls to finish or add credit. Agents earn credit from their token's trading fees. - Retry
- Fix first
- Cost
- Free
{ "error": { "message": "Balance $0.004210 ($0.001000 held by calls running or being priced) is below this request's reserve of $0.019660. Lower the output limit or add credit.", "type": "billing_error", "code": "insufficient_balance", "param": null }}{ "type": "error", "error": { "type": "billing_error", "message": "Balance $0.004210 ($0.001000 held by calls running or being priced) is below this request's reserve of $0.019660. Lower the output limit or add credit.", "code": "insufficient_balance" }}402key_limitKey's spending limit reachedFix first
- When
- The key's own daily, weekly or monthly limit can't cover this call's reserve. The agent's other keys keep working.
- What to do
- Wait for the reset named in the message, lower
max_tokensor raise the limit on the keys page. - Retry
- Fix first
- Cost
- Free
{ "error": { "message": "This key's limit of $5.000000 a day is used up: $4.990000 spent and $0.000000 held by calls running or being priced, and this request reserves $0.019660. It resets at 2026-10-01T00:00:00.000Z. Lower the output limit, or ask the agent's owner to raise the key's limit.", "type": "billing_error", "code": "key_limit", "param": null }}{ "type": "error", "error": { "type": "billing_error", "message": "This key's limit of $5.000000 a day is used up: $4.990000 spent and $0.000000 held by calls running or being priced, and this request reserves $0.019660. It resets at 2026-10-01T00:00:00.000Z. Lower the output limit, or ask the agent's owner to raise the key's limit.", "code": "key_limit" }}429rate_limitedToo fastRetry after the wait
- When
- The agent started more than 60 calls at once, or more than 600 a minute.
- What to do
- Wait for Retry-After-Ms and send again. Official OpenAI and Anthropic SDKs do this for you.
- Retry
- Retry after the wait in Retry-After-Ms
- Cost
- Free
{ "error": { "message": "This agent is starting calls faster than 600 a minute. Retry in 1 second.", "type": "rate_limit_error", "code": "rate_limited", "param": null }}{ "type": "error", "error": { "type": "rate_limit_error", "message": "This agent is starting calls faster than 600 a minute. Retry in 1 second.", "code": "rate_limited" }}429too_many_running_calls8 calls already runningRetry after the wait
- When
- The agent has 8 calls running. A slot frees the moment one ends.
- What to do
- Queue calls on your side, or wait the 2 to 8 seconds the headers name.
- Retry
- Retry after the wait in Retry-After-Ms
- Cost
- Free
{ "error": { "message": "This agent already has 8 calls running. Retry when one finishes.", "type": "rate_limit_error", "code": "too_many_running_calls", "param": null }}{ "type": "error", "error": { "type": "rate_limit_error", "message": "This agent already has 8 calls running. Retry when one finishes.", "code": "too_many_running_calls" }}503/529model_busyModel busyRetry after the wait
- When
- Every provider of this model is at capacity. On /messages it is 529
overloaded_error, so Claude Code moves to its fallback model. - What to do
- Wait for Retry-After-Ms, or use another model. Listing fallbacks in
modelslets a call run on the next one. - Retry
- Retry after the wait in Retry-After-Ms
- Cost
- Free
{ "error": { "message": "The model `openai/gpt-6.1` is busy at its providers right now. Retry in 12 seconds.", "type": "server_error", "code": "model_busy", "param": null }}{ "type": "error", "error": { "type": "overloaded_error", "message": "The model `openai/gpt-6.1` is busy at its providers right now. Retry in 12 seconds.", "code": "model_busy" }}503gateway_busyGateway busyRetry after the wait
- When
- The gateway is at its capacity for a moment.
- What to do
- Wait for Retry-After-Ms and send again.
- Retry
- Retry after the wait in Retry-After-Ms
- Cost
- Free
{ "error": { "message": "The gateway is busy right now. Retry in 8 seconds.", "type": "server_error", "code": "gateway_busy", "param": null }}{ "type": "error", "error": { "type": "overloaded_error", "message": "The gateway is busy right now. Retry in 8 seconds.", "code": "gateway_busy" }}503gateway_capacityOut of capacity brieflyRetry after the wait
- When
- The gateway is topping up its capacity. bInference is alerted at once.
- What to do
- Wait for Retry-After-Ms and send again.
- Retry
- Retry after the wait in Retry-After-Ms
- Cost
- Free
{ "error": { "message": "The gateway is briefly out of capacity. Try again soon.", "type": "server_error", "code": "gateway_capacity", "param": null }}{ "type": "error", "error": { "type": "overloaded_error", "message": "The gateway is briefly out of capacity. Try again soon.", "code": "gateway_capacity" }}500gateway_errorGateway failedRetry
- When
- A fault in the gateway after the reserve. The reserve is released and the call is free.
- What to do
- Send it again. If it keeps failing, contact bInference with the
x-binference-call-idheader. - Retry
- Retry
- Cost
- Free
{ "error": { "message": "The gateway failed on this call. Retry it.", "type": "server_error", "code": "gateway_error", "param": null }}{ "type": "error", "error": { "type": "api_error", "message": "The gateway failed on this call. Retry it.", "code": "gateway_error" }}504upstream_timeoutNo answer in timeRetry
- When
- The model sent nothing before the time limit: 30 minutes streamed, 13 minutes not streamed.
- What to do
- Stream long calls, lower
max_tokensor use a faster model. - Retry
- Retry
- Cost
- Only what the provider recorded, usually nothing
{ "error": { "message": "The model did not answer in time.", "type": "server_error", "code": "upstream_timeout", "param": null }}{ "type": "error", "error": { "type": "api_error", "message": "The model did not answer in time.", "code": "upstream_timeout" }}502upstream_unreachableProvider unreachableRetry
- When
- The connection to the model's provider failed.
- What to do
- Send it again.
- Retry
- Retry
- Cost
- Free
{ "error": { "message": "Could not reach the model provider. Try again.", "type": "server_error", "code": "upstream_unreachable", "param": null }}{ "type": "error", "error": { "type": "api_error", "message": "Could not reach the model provider. Try again.", "code": "upstream_unreachable" }}502provider_errorProvider failedRetry
- When
- The model's provider failed, or refused this one call.
- What to do
- Send it again, or use another model.
- Retry
- Retry
- Cost
- Only what the provider recorded, usually nothing
{ "error": { "message": "The model's provider failed to answer. Try again in a moment.", "type": "server_error", "code": "provider_error", "param": null }}{ "type": "error", "error": { "type": "api_error", "message": "The model's provider failed to answer. Try again in a moment.", "code": "provider_error" }}404no_providerNo provider for this requestFix first
- When
- No provider of the model supports every option in the request, such as tools, images or a long context.
- What to do
- Drop an option, or pick a model that lists it.
- Retry
- Fix first
- Cost
- Free
{ "error": { "message": "No provider can serve this request with this model. Try another model, or fewer options such as tools or images.", "type": "not_found_error", "code": "no_provider", "param": null }}{ "type": "error", "error": { "type": "not_found_error", "message": "No provider can serve this request with this model. Try another model, or fewer options such as tools or images.", "code": "no_provider" }}403provider_refusedProvider refusedFix first
- When
- The provider's own policy refused the request.
- What to do
- Change the request, or use another model.
- Retry
- Fix first
- Cost
- Free
{ "error": { "message": "The model's provider refused this request for this model.", "type": "permission_error", "code": "provider_refused", "param": null }}{ "type": "error", "error": { "type": "permission_error", "message": "The model's provider refused this request for this model.", "code": "provider_refused" }}422unprocessableProvider couldn't process itFix first
- When
- The provider accepted the request but couldn't run it.
- What to do
- Check the parameters against the model, or use another model.
- Retry
- Fix first
- Cost
- Free
{ "error": { "message": "The model's provider could not process this request.", "type": "server_error", "code": "unprocessable", "param": null }}{ "type": "error", "error": { "type": "api_error", "message": "The model's provider could not process this request.", "code": "unprocessable" }}408timeoutProvider timed outRetry
- When
- The provider gave up on the call (408 or 504).
- What to do
- Send it again, stream it or lower
max_tokens. - Retry
- Retry
- Cost
- Only what the provider recorded, usually nothing
{ "error": { "message": "The model took too long to answer. Try again.", "type": "server_error", "code": "timeout", "param": null }}{ "type": "error", "error": { "type": "api_error", "message": "The model took too long to answer. Try again.", "code": "timeout" }}502upstream_errorOther provider errorRetry
- When
- Any other error from the provider.
- What to do
- Send it again shortly.
- Retry
- Retry
- Cost
- Only what the provider recorded, usually nothing
{ "error": { "message": "The request could not be completed. Try again in a moment.", "type": "server_error", "code": "upstream_error", "param": null }}{ "type": "error", "error": { "type": "api_error", "message": "The request could not be completed. Try again in a moment.", "code": "upstream_error" }}The short version
| Status | Means | Retry? |
|---|---|---|
400 | The request can't be served as sent | No. Fix it first |
401 | The key is missing, wrong or revoked | No |
402 | The balance or the key's limit can't cover the reserve | No. Lower max_tokens or add credit |
403 | The agent isn't active | No |
404 | Unknown path, or no provider for this request | No |
413 | The body is over 4 MB | No. Send files as URLs |
429 | Too many running calls, or too fast | Yes, after Retry-After-Ms |
500 | The gateway failed. The call is free | Yes |
502 504 | The model's provider failed | Yes |
503 529 | Busy for a moment | Yes, after Retry-After-Ms |
Errors inside a stream
Once an answer has started streaming, an error can't change the status any more. It arrives as the format's own error event, so SDKs raise it instead of hanging:
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"error"}],"error":{"message":"The model took too long to answer. Try again.","type":"server_error","code":"timeout"}}The call is charged for what was written before the error.