API overview
One base URL, one key, three formats. Everything you need before the first call.
Base URL
https://binference.io/api/v1For the Anthropic SDKs and Claude Code, use https://binference.io/api: they add /v1 themselves.
Endpoints
| Method | Path | Format | Key |
|---|---|---|---|
POST | /chat/completions | OpenAI Chat Completions | Yes |
POST | /responses | OpenAI Responses | Yes |
POST | /messages | Anthropic Messages | Yes |
GET | /models | Every model and its price | No |
GET | /balance | What the key can spend | Yes |
Any other path under /api/v1, like embeddings, answers 404 not_found in JSON.
Authentication
Send your binf_ key as a bearer token or as x-api-key:
Authorization: Bearer binf_...A missing, wrong or revoked key answers 401 invalid_api_key. See API keys.
Requests
- JSON bodies, up to 4 MB. Send images and files as URLs, not inline data.
- Bodies pass through as sent. Every field of each format works, beyond the ones these pages list.
- Stateless. Nothing is stored between calls. Send the conversation every time.
- From anywhere. The API accepts calls from any origin, so browser tools work. Keys in web pages can be read by visitors: keep them on a server.
Answers
Answers follow each format's own shape, with three additions:
| Field or header | What it is |
|---|---|
usage.cost | What the call was charged, in dollars. On Chat Completions and Responses. |
x-binference-call-id | The call's id on your Calls page. Quote it when you contact us. |
x-generation-id | The provider's id for the answer. |
Errors
Errors use each format's own shape, with a stable code to branch on:
{
"error": {
"message": "This agent is starting calls faster than 600 a minute. Retry in 1 second.",
"type": "rate_limit_error",
"code": "rate_limited",
"param": null
}
}See every code in Errors, and the limits in Rate limits.