Messages

The Anthropic Messages format, for Claude Code and every model.

POSTTry it

Anthropic's format, used by Claude Code and the Anthropic SDKs. It works with every listed model, not only Claude.

max_tokens is required, as in Anthropic's own API. A busy model answers 529 overloaded_error, so Claude Code moves to its fallback model. Messages answers carry no cost: find each call's charge on your Calls page.

Send your key as Authorization: Bearer binf_... or x-api-key.

Request body

The fields most calls use. Every other field of the Anthropic format is passed on as sent.

  • modelstringRequired

    A model id from GET /models. Claude clients' own names work too: claude-sonnet-5-5, dated names and the [1m] suffix map to the listed model.

  • messagesobject[]Required

    The conversation so far, alternating user and assistant.

  • max_tokensintegerRequired

    The longest answer, in tokens. The call reserves this much output while it runs.

  • systemstring | object[]

    The system prompt.

  • streamboolean

    Send the answer as it is written, as server-sent events. Recommended for anything longer than a sentence: streams may run 30 minutes, other calls 13.

  • temperaturenumber

    Randomness, 0 to 2. Lower is more predictable.

  • toolsobject[]

    Functions the model may call, in the format's own shape. Your code runs them and sends the results back. Web search and web fetch are also served.

  • thinkingobject

    { "type": "enabled", "budget_tokens": 2048 } for extended thinking. Thinking is billed as output.

Headers

  • anthropic-versionstring

    Passed on as sent. The Anthropic SDKs set it.

  • anthropic-betastring

    Passed on as sent, for beta features.

Response

  • idstring

    The message's id.

  • contentobject[]

    Blocks: text, thinking and tool_use.

  • stop_reasonstring

    Why the answer ended.

  • usageobject

    input_tokens and output_tokens. The Messages format carries no cost: find each call's charge on your Calls page within a minute.

Response headers

  • x-binference-call-idstring

    This call's id on your Calls page. Quote it when you contact us.

  • x-generation-idstring

    The model provider's id for the answer.

  • retry-afterinteger

    On a 429 or 503: whole seconds to wait before sending again.

  • retry-after-msinteger

    The same wait in milliseconds. The OpenAI and Anthropic SDKs read it.

Request
curl https://binference.io/api/v1/messages \  -H "Authorization: Bearer $BINF_API_KEY" \  -H "anthropic-version: 2023-06-01" \  -H "Content-Type: application/json" \  -d '{    "model": "anthropic/claude-sonnet-5.5",    "max_tokens": 256,    "system": "You are a concise trading assistant.",    "messages": [      {        "role": "user",        "content": "Summarize the last 24 hours of trades in one line."      }    ]  }'
Response
{  "id": "msg_01XFDUDYJgAACzvnptvVoYEL",  "type": "message",  "role": "assistant",  "model": "anthropic/claude-sonnet-5.5",  "content": [    {      "type": "text",      "text": "12 trades today: 9 wins, net +3.4 BNB, biggest move on $NOVA."    }  ],  "stop_reason": "end_turn",  "usage": {    "input_tokens": 27,    "output_tokens": 22  }}
Streamed, with stream: true
event: message_startdata: {"type":"message_start","message":{"id":"msg_01XFDUDYJgAACzvnptvVoYEL","role":"assistant","usage":{"input_tokens":27,"output_tokens":1}}}event: content_block_deltadata: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"12 trades today"}}event: message_deltadata: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":22}}event: message_stopdata: {"type":"message_stop"}