Chat Completions

The OpenAI chat format, for every model.

POSTTry it

Send a conversation, get the next message. It's the format most tools speak: the OpenAI SDKs, the Vercel AI SDK, LangChain, OpenClaw and Hermes.

Stream long answers with stream: true, and add stream_options: { "include_usage": true } to get the cost in the last chunk. See Streaming.

Send your key as Authorization: Bearer binf_... or x-api-key.

Request body

The fields most calls use. Every other field of the OpenAI format is passed on as sent.

  • modelstringRequired

    A model id from GET /models, such as anthropic/claude-sonnet-5.5. Add :online for web search, or :nitro, :floor or :exacto to steer which provider runs it.

  • messagesobject[]Required

    The conversation so far, oldest first.

    • role"system" | "user" | "assistant" | "tool"Required

      Who wrote the message.

    • contentstring | object[]Required

      Text, or parts: text and image URLs. Send images and files as URLs, not inline data: a request body is at most 4 MB.

  • max_tokensinteger

    The longest answer, in tokens. Set it: a call reserves its longest possible answer while it runs, and without a limit that is the model's whole output. max_completion_tokens works too.

  • streamboolean

    Send the answer as it is written, as server-sent events. Recommended for anything longer than a sentence: streams may run 30 minutes, other calls 13.

  • stream_optionsobject

    { "include_usage": true } adds a last chunk with usage, including what the call cost.

  • temperaturenumber

    Randomness, 0 to 2. Lower is more predictable.

  • toolsobject[]

    Functions the model may call, in the format's own shape. Your code runs them and sends the results back. Web search and web fetch are also served.

  • tool_choicestring | object

    "auto", "none", "required" or one function by name.

  • response_formatobject

    Structured output: { "type": "json_schema", "json_schema": { ... } } makes the answer match your schema.

  • reasoningobject

    For reasoning models: { "effort": "low" | "medium" | "high" }. Reasoning tokens are billed as output.

  • modelsstring[]

    Fallback models, tried in order when the first is busy or down. The reserve covers the dearest of them.

Response

  • idstring

    The answer's id.

  • choicesobject[]

    The answer: message.content, any message.tool_calls, and finish_reason.

  • usageobject

    Tokens and what the call cost.

    • prompt_tokensinteger

      Tokens read.

    • completion_tokensinteger

      Tokens written, reasoning included.

    • costnumber

      What this call is charged, in dollars. Exactly what leaves the balance.

Response headers

  • x-binference-call-idstring

    This call's id on your Calls page. Quote it when you contact us.

  • x-generation-idstring

    The model provider's id for the answer.

  • retry-afterinteger

    On a 429 or 503: whole seconds to wait before sending again.

  • retry-after-msinteger

    The same wait in milliseconds. The OpenAI and Anthropic SDKs read it.

Request
curl https://binference.io/api/v1/chat/completions \  -H "Authorization: Bearer $BINF_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "model": "anthropic/claude-sonnet-5.5",    "messages": [      {        "role": "system",        "content": "You are a concise trading assistant."      },      {        "role": "user",        "content": "Summarize the last 24 hours of trades in one line."      }    ],    "max_tokens": 256  }'
Response
{  "id": "gen-1790755652-kQ3hTz",  "object": "chat.completion",  "created": 1790755652,  "model": "anthropic/claude-sonnet-5.5",  "choices": [    {      "index": 0,      "message": {        "role": "assistant",        "content": "12 trades today: 9 wins, net +3.4 BNB, biggest move on $NOVA."      },      "finish_reason": "stop"    }  ],  "usage": {    "prompt_tokens": 31,    "completion_tokens": 22,    "total_tokens": 53,    "cost": 0.000338  }}
Streamed, with stream: true
data: {"id":"gen-1790755652-kQ3hTz","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":"12 trades"}}]}data: {"id":"gen-1790755652-kQ3hTz","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" today: 9 wins"}}]}data: {"id":"gen-1790755652-kQ3hTz","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":31,"completion_tokens":22,"total_tokens":53,"cost":0.000338}}data: [DONE]