Chat Completions
The OpenAI chat format, for every model.
Send a conversation, get the next message. It's the format most tools speak: the OpenAI SDKs, the Vercel AI SDK, LangChain, OpenClaw and Hermes.
Stream long answers with stream: true, and add stream_options: { "include_usage": true } to get the cost in the last chunk. See Streaming.
Authorization: Bearer binf_... or x-api-key.Request body
The fields most calls use. Every other field of the OpenAI format is passed on as sent.
modelstringRequiredA model id from GET /models, such as
anthropic/claude-sonnet-5.5. Add:onlinefor web search, or:nitro,:flooror:exactoto steer which provider runs it.messagesobject[]RequiredThe conversation so far, oldest first.
role"system" | "user" | "assistant" | "tool"RequiredWho wrote the message.
contentstring | object[]RequiredText, or parts: text and image URLs. Send images and files as URLs, not inline data: a request body is at most 4 MB.
max_tokensintegerThe longest answer, in tokens. Set it: a call reserves its longest possible answer while it runs, and without a limit that is the model's whole output.
max_completion_tokensworks too.streambooleanSend the answer as it is written, as server-sent events. Recommended for anything longer than a sentence: streams may run 30 minutes, other calls 13.
stream_optionsobject{ "include_usage": true }adds a last chunk withusage, including what the call cost.temperaturenumberRandomness, 0 to 2. Lower is more predictable.
toolsobject[]Functions the model may call, in the format's own shape. Your code runs them and sends the results back. Web search and web fetch are also served.
tool_choicestring | object"auto","none","required"or one function by name.response_formatobjectStructured output:
{ "type": "json_schema", "json_schema": { ... } }makes the answer match your schema.reasoningobjectFor reasoning models:
{ "effort": "low" | "medium" | "high" }. Reasoning tokens are billed as output.modelsstring[]Fallback models, tried in order when the first is busy or down. The reserve covers the dearest of them.
Response
idstringThe answer's id.
choicesobject[]The answer:
message.content, anymessage.tool_calls, andfinish_reason.usageobjectTokens and what the call cost.
prompt_tokensintegerTokens read.
completion_tokensintegerTokens written, reasoning included.
costnumberWhat this call is charged, in dollars. Exactly what leaves the balance.
Response headers
x-binference-call-idstringThis call's id on your Calls page. Quote it when you contact us.
x-generation-idstringThe model provider's id for the answer.
retry-afterintegerOn a 429 or 503: whole seconds to wait before sending again.
retry-after-msintegerThe same wait in milliseconds. The OpenAI and Anthropic SDKs read it.
curl https://binference.io/api/v1/chat/completions \ -H "Authorization: Bearer $BINF_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "anthropic/claude-sonnet-5.5", "messages": [ { "role": "system", "content": "You are a concise trading assistant." }, { "role": "user", "content": "Summarize the last 24 hours of trades in one line." } ], "max_tokens": 256 }'{ "id": "gen-1790755652-kQ3hTz", "object": "chat.completion", "created": 1790755652, "model": "anthropic/claude-sonnet-5.5", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "12 trades today: 9 wins, net +3.4 BNB, biggest move on $NOVA." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 31, "completion_tokens": 22, "total_tokens": 53, "cost": 0.000338 }}data: {"id":"gen-1790755652-kQ3hTz","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":"12 trades"}}]}data: {"id":"gen-1790755652-kQ3hTz","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" today: 9 wins"}}]}data: {"id":"gen-1790755652-kQ3hTz","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":31,"completion_tokens":22,"total_tokens":53,"cost":0.000338}}data: [DONE]