Responses
The OpenAI Responses format, which Codex speaks.
The newer OpenAI format, with input and instructions in place of messages. Codex uses it.
It is stateless here: nothing is stored, so previous_response_id and saved prompts aren't served. Send the conversation as input items each time.
Authorization: Bearer binf_... or x-api-key.Request body
The fields most calls use. Every other field of the OpenAI format is passed on as sent.
modelstringRequiredA model id from GET /models, such as
anthropic/claude-sonnet-5.5. Add:onlinefor web search, or:nitro,:flooror:exactoto steer which provider runs it.inputstring | object[]RequiredA prompt, or the conversation as input items.
instructionsstringThe system prompt.
max_output_tokensintegerThe longest answer, in tokens. Set it: the call reserves its longest possible answer while it runs.
streambooleanSend the answer as it is written, as server-sent events. Recommended for anything longer than a sentence: streams may run 30 minutes, other calls 13.
temperaturenumberRandomness, 0 to 2. Lower is more predictable.
toolsobject[]Functions the model may call, in the format's own shape. Your code runs them and sends the results back. Web search and web fetch are also served.
max_tool_callsintegerThe most tool steps in one call. Web search loops reserve for every step, so a lower number reserves less.
reasoningobjectFor reasoning models:
{ "effort": "low" | "medium" | "high" }.modelsstring[]Fallback models, tried in order when the first is busy or down. The reserve covers the dearest of them.
Response
idstringThe response's id.
outputobject[]Output items: messages, reasoning and tool calls.
output_textjoins the text.usageobjectinput_tokens,output_tokensandcost: what this call is charged, in dollars.
Response headers
x-binference-call-idstringThis call's id on your Calls page. Quote it when you contact us.
x-generation-idstringThe model provider's id for the answer.
retry-afterintegerOn a 429 or 503: whole seconds to wait before sending again.
retry-after-msintegerThe same wait in milliseconds. The OpenAI and Anthropic SDKs read it.
curl https://binference.io/api/v1/responses \ -H "Authorization: Bearer $BINF_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "openai/gpt-6.1-sol", "instructions": "You are a concise trading assistant.", "input": "Summarize the last 24 hours of trades in one line.", "max_output_tokens": 256 }'{ "id": "resp_7Yq2c1", "object": "response", "status": "completed", "model": "openai/gpt-6.1-sol", "output": [ { "type": "message", "role": "assistant", "content": [ { "type": "output_text", "text": "12 trades today: 9 wins, net +3.4 BNB, biggest move on $NOVA." } ] } ], "usage": { "input_tokens": 29, "output_tokens": 21, "total_tokens": 50, "cost": 0.000322 }}event: response.createddata: {"type":"response.created","response":{"id":"resp_7Yq2c1","status":"in_progress"}}event: response.output_text.deltadata: {"type":"response.output_text.delta","delta":"12 trades today"}event: response.completeddata: {"type":"response.completed","response":{"id":"resp_7Yq2c1","status":"completed","usage":{"input_tokens":29,"output_tokens":21,"total_tokens":50,"cost":0.000322}}}