API overview

One base URL, one key, three formats. Everything you need before the first call.

Base URL

https://binference.io/api/v1

For the Anthropic SDKs and Claude Code, use https://binference.io/api: they add /v1 themselves.

Endpoints

MethodPathFormatKey
POST/chat/completionsOpenAI Chat CompletionsYes
POST/responsesOpenAI ResponsesYes
POST/messagesAnthropic MessagesYes
GET/modelsEvery model and its priceNo
GET/balanceWhat the key can spendYes

Any other path under /api/v1, like embeddings, answers 404 not_found in JSON.

Authentication

Send your binf_ key as a bearer token or as x-api-key:

Authorization: Bearer binf_...

A missing, wrong or revoked key answers 401 invalid_api_key. See API keys.

Requests

  • JSON bodies, up to 4 MB. Send images and files as URLs, not inline data.
  • Bodies pass through as sent. Every field of each format works, beyond the ones these pages list.
  • Stateless. Nothing is stored between calls. Send the conversation every time.
  • From anywhere. The API accepts calls from any origin, so browser tools work. Keys in web pages can be read by visitors: keep them on a server.

Answers

Answers follow each format's own shape, with three additions:

Field or headerWhat it is
usage.costWhat the call was charged, in dollars. On Chat Completions and Responses.
x-binference-call-idThe call's id on your Calls page. Quote it when you contact us.
x-generation-idThe provider's id for the answer.

Errors

Errors use each format's own shape, with a stable code to branch on:

{
  "error": {
    "message": "This agent is starting calls faster than 600 a minute. Retry in 1 second.",
    "type": "rate_limit_error",
    "code": "rate_limited",
    "param": null
  }
}

See every code in Errors, and the limits in Rate limits.

On this page