Messages
The Anthropic Messages format, for Claude Code and every model.
Anthropic's format, used by Claude Code and the Anthropic SDKs. It works with every listed model, not only Claude.
max_tokens is required, as in Anthropic's own API. A busy model answers 529 overloaded_error, so Claude Code moves to its fallback model. Messages answers carry no cost: find each call's charge on your Calls page.
Authorization: Bearer binf_... or x-api-key.Request body
The fields most calls use. Every other field of the Anthropic format is passed on as sent.
modelstringRequiredA model id from GET /models. Claude clients' own names work too:
claude-sonnet-5-5, dated names and the[1m]suffix map to the listed model.messagesobject[]RequiredThe conversation so far, alternating
userandassistant.max_tokensintegerRequiredThe longest answer, in tokens. The call reserves this much output while it runs.
systemstring | object[]The system prompt.
streambooleanSend the answer as it is written, as server-sent events. Recommended for anything longer than a sentence: streams may run 30 minutes, other calls 13.
temperaturenumberRandomness, 0 to 2. Lower is more predictable.
toolsobject[]Functions the model may call, in the format's own shape. Your code runs them and sends the results back. Web search and web fetch are also served.
thinkingobject{ "type": "enabled", "budget_tokens": 2048 }for extended thinking. Thinking is billed as output.
Headers
anthropic-versionstringPassed on as sent. The Anthropic SDKs set it.
anthropic-betastringPassed on as sent, for beta features.
Response
idstringThe message's id.
contentobject[]Blocks:
text,thinkingandtool_use.stop_reasonstringWhy the answer ended.
usageobjectinput_tokensandoutput_tokens. The Messages format carries no cost: find each call's charge on your Calls page within a minute.
Response headers
x-binference-call-idstringThis call's id on your Calls page. Quote it when you contact us.
x-generation-idstringThe model provider's id for the answer.
retry-afterintegerOn a 429 or 503: whole seconds to wait before sending again.
retry-after-msintegerThe same wait in milliseconds. The OpenAI and Anthropic SDKs read it.
curl https://binference.io/api/v1/messages \ -H "Authorization: Bearer $BINF_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "anthropic/claude-sonnet-5.5", "max_tokens": 256, "system": "You are a concise trading assistant.", "messages": [ { "role": "user", "content": "Summarize the last 24 hours of trades in one line." } ] }'{ "id": "msg_01XFDUDYJgAACzvnptvVoYEL", "type": "message", "role": "assistant", "model": "anthropic/claude-sonnet-5.5", "content": [ { "type": "text", "text": "12 trades today: 9 wins, net +3.4 BNB, biggest move on $NOVA." } ], "stop_reason": "end_turn", "usage": { "input_tokens": 27, "output_tokens": 22 }}event: message_startdata: {"type":"message_start","message":{"id":"msg_01XFDUDYJgAACzvnptvVoYEL","role":"assistant","usage":{"input_tokens":27,"output_tokens":1}}}event: content_block_deltadata: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"12 trades today"}}event: message_deltadata: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":22}}event: message_stopdata: {"type":"message_stop"}