# Agentic Wallet (https://docs.binference.io/agent-os/agentic-wallet)
Let your agent swap, set limit orders and send tokens on chain, within rules you set in the Binance App.
The Binance Agentic Wallet is an on-chain wallet your agent drives in plain language: "Swap 0.1 BNB for USDT." It's separate from the MCP Server, which trades on the Binance exchange. Use either or both.
It installs as a skill, so it works in Claude Code, Codex and OpenClaw, each thinking with bInference.
## What it can do [#what-it-can-do]
| Ask | Where |
| ---------------------------------------------------------------------- | --------------------------------------------- |
| "How much BNB and USDT do I have? How much of my daily limit is left?" | All chains |
| "Swap 0.1 BNB for USDT" | All chains |
| "Buy BNB with 100 USDT when the price drops to $500" | BNB Chain and Solana |
| "Send 10 USDT to 0x1234...5678" | All chains, to addresses in your address book |
| "Buy Up in the most recent BTC 15m price market" | Prediction markets |
| "Deposit 100 USDT to Venus" | DeFi on supported protocols |
Chains: BNB Chain, Ethereum, Base and Solana.
## Set it up [#set-it-up]
### Make an MPC Wallet [#make-an-mpc-wallet]
In the Binance App, sign in and create an MPC Wallet. You need it before an Agentic Wallet can exist. You also need Node.js 18 or newer.
### Install the skill [#install-the-skill]
```bash
npx skills add binance/binance-skills-hub/skills/binance-web3/binance-agentic-wallet
```
The skill installs Binance's `baw` command line tool itself the first time it's used.
### Sign in from your agent [#sign-in-from-your-agent]
Say to your agent:
```txt
Sign in to Binance Agentic Wallet
```
It answers with a sign-in link. On a phone, the link opens the Binance App. On a computer, scan the page's QR code with the Binance App and confirm. The first time, the App walks you through creating the Agentic Wallet.
### Set its rules in the Binance App [#set-its-rules-in-the-binance-app]
Rules can only be changed in the App: the agent can read them, never change them. Binance recommends starting with:
* a small daily limit, such as $50
* tradable tokens set to **Limited**
* high-risk transactions set to **Require App confirmation**
Find them in **Binance App → Agentic Wallet management page → Settings** (top right).
## Your first swap [#your-first-swap]
1. "Is my wallet connected? What's my address?"
2. "How much BNB and USDT do I have?"
3. "How much USDT would I get for 0.01 BNB? Don't execute."
4. "Go ahead and execute." The agent asks you to confirm, then sends the swap and returns an order id.
Keep a little BNB in the wallet for gas.
## Good to know [#good-to-know]
* **Your key stays yours.** The wallet uses MPC: the private key is never whole on any device or server, and the agent never holds it.
* **Sign out to stop it.** Tell your agent "Sign out" and it loses access at once.
* **Skills add tokens to every call.** The skill's instructions are read by the model, so they count as input on bInference. See [What it costs](/agent-os/costs).
# Claude Code (https://docs.binference.io/agent-os/claude-code)
Run Claude Code on bInference and connect it to Binance, from key to first trade.
### Get a bInference key [#get-a-binference-key]
On [binference.io](https://binference.io/account/keys), sign in with your wallet and press **Create key**. Turn on **Limit spending** and set a daily limit: it caps what this agent can spend on AI, however long it runs.
### Point Claude Code at bInference [#point-claude-code-at-binference]
Add these to `~/.claude/settings.json`, so every Claude Code session uses them:
```json title="~/.claude/settings.json"
{
"env": {
"ANTHROPIC_BASE_URL": "https://binference.io/api",
"ANTHROPIC_AUTH_TOKEN": "binf_...",
"ANTHROPIC_API_KEY": ""
}
}
```
`ANTHROPIC_API_KEY` stays empty so Claude Code never falls back to another key. To pick the model, add `"ANTHROPIC_MODEL": "anthropic/claude-sonnet-5.5"`. Any model on [Models](/models) that handles tools well works.
### Add the Binance MCP Server [#add-the-binance-mcp-server]
In Terminal (PowerShell on Windows), run Binance's command exactly as written:
```bash
claude mcp add binance-mcp-server --transport http https://agent.binance.com/mcp/agentic
```
Then start Claude Code, open the `/mcp` menu, select **binance-mcp-server** and authenticate. Binance opens its own sign-in and asks which scopes to grant. Start with the fewest:
| Scope | Lets the agent |
| ----------- | -------------------------------------------------------------------------------------------- |
| Market data | Read prices, order books, candles and funding. No sign-in needed |
| Account | Read the Agentic sub-account's balances and positions, and optionally view your main account |
| Trade | Trade Spot, Margin, Convert and Futures, as far as your account allows |
| Transfer | Move funds between wallets inside the Agentic sub-account only |
There is no withdrawal scope: the agent can never send funds out.
Don't ask Claude to install the server for you, and don't open the address in a
browser. Binance asks you to run the command above yourself.
### Check both connections [#check-both-connections]
Ask Claude Code:
```txt
Use the Binance MCP Server to show the current BTCUSDT price and 24-hour change.
```
The answer should show the Binance tool being used and live BTCUSDT data. That one question used both sides: the model ran on bInference, the price came from Binance. The call shows up on your [Calls page](https://binference.io/account/calls) within a minute.
### Fund the Agentic sub-account [#fund-the-agentic-sub-account]
Your agent trades in its own sub-account, created the first time you authorize. It starts empty, and only you can fund it:
1. On binance.com, go to **Profile → Dashboard → Sub-account → Asset Management**.
2. Press **Transfer** and move assets from your main account to the Agentic sub-account. It's instant and free.
Put in only what you're willing to let the agent trade.
## Your first session [#your-first-session]
| You ask | What happens |
| --------------------------------------------------------- | ---------------------------------------------------------------------- |
| "What's BTCUSDT trading at, and how's it moved over 24h?" | Reads live data at once. Nothing to confirm |
| "What do I have in my Agentic account?" | Reads the sub-account's balances across wallets |
| "Buy $100 of BNB at market on spot." | Restates the order and waits for your yes. This step spends real funds |
| "Did that fill? What's my BNB balance now?" | Checks the order and the new balance |
Every order, cancel and transfer asks for your confirmation first. Reads run straight away.
## Add on-chain trading [#add-on-chain-trading]
To let the same Claude Code session swap tokens on BNB Chain, Ethereum, Base or Solana, add the Agentic Wallet skill. See [Agentic Wallet](/agent-os/agentic-wallet).
## Good to know [#good-to-know]
* **Long sessions cost more per turn.** Claude Code sends the whole conversation, plus Binance's tool list, on every turn. Run `/compact` when a session grows long. See [What it costs](/agent-os/costs).
* **Busy models:** a busy model answers `529 overloaded_error` and Claude Code moves to its fallback model on its own.
* **Big pasted images** can push a request past 4 MB (`413`). Run `/compact` or `/clear`.
* **Two different "insufficient balance" errors.** A trade that fails for funds means the Binance sub-account is empty. A `402 insufficient_balance` from bInference means the AI budget is. See [Troubleshooting](/agent-os/troubleshooting).
# Codex CLI (https://docs.binference.io/agent-os/codex)
Run Codex on bInference and connect it to Binance. Works for the Codex and ChatGPT desktop apps too.
### Install Codex [#install-codex]
Skip this if `codex` already runs in your terminal.
macOS and Linux
Windows
```bash
curl -fsSL https://chatgpt.com/codex/install.sh | sh
```
```powershell
powershell -ExecutionPolicy ByPass -c "irm https://chatgpt.com/codex/install.ps1 | iex"
```
### Get a bInference key [#get-a-binference-key]
On [binference.io](https://binference.io/account/keys), press **Create key** and set a daily spending limit. Keep the key in your environment:
```bash
export BINF_API_KEY="binf_..."
```
### Point Codex at bInference [#point-codex-at-binference]
Codex speaks the Responses format. Add bInference as a provider:
```toml title="~/.codex/config.toml"
model_provider = "binference"
model = "openai/gpt-6.1-sol"
[model_providers.binference]
name = "bInference"
base_url = "https://binference.io/api/v1"
env_key = "BINF_API_KEY"
wire_api = "responses"
```
Any model on [Models](/models) works as `model`.
### Add the Binance MCP Server [#add-the-binance-mcp-server]
Run Binance's command exactly as written, then authenticate when the browser opens:
```bash
codex mcp add binance-mcp-server --url https://agent.binance.com/mcp/agentic --oauth-client-id codex
```
Grant the fewest scopes you need: market data, account, trade or transfer. There is no withdrawal scope. The Codex and ChatGPT desktop apps pick up the same connection.
### Check it, then fund the sub-account [#check-it-then-fund-the-sub-account]
Ask Codex:
```txt
Use the Binance MCP Server to show the current BTCUSDT price and 24-hour change.
```
Then move trading funds in on binance.com: **Profile → Dashboard → Sub-account → Asset Management → Transfer**. Only you can fund the Agentic sub-account, and it starts empty.
## Good to know [#good-to-know]
* **Retries are built in.** Codex retries `503` answers, so a busy model or a short wait on our side resolves on its own.
* **Every order waits for you.** The agent restates each order, cancel or transfer and waits for your yes.
* **Two limits keep it safe:** Binance's scopes and sub-account on one side, the key's spending limit here on the other. See [Safety and limits](/agent-os/safety).
# What it costs (https://docs.binference.io/agent-os/costs)
What an Agent OS session spends on AI, and five ways to keep it low.
Trading through Agent OS is between you and Binance. What your agent spends on bInference is its model calls: tokens in and out, at the prices on [Models](/models).
## Where the tokens go [#where-the-tokens-go]
Every turn of a session sends the model:
* **the conversation so far,** which grows each turn
* **Binance's tool list,** so the model knows what it can do
* **tool results,** like an order book or a balance, read as input
* **any skills** the question calls for
The model's answer, and any reasoning, is output.
## An example session [#an-example-session]
Twenty turns with Claude Sonnet 5.5 at $2.40 per million tokens in and $12 out, averaging 25,000 tokens in and 800 out per turn:
| | Tokens | Cost |
| ----------- | ------- | --------- |
| In | 500,000 | $1.20 |
| Out | 16,000 | $0.19 |
| **Session** | | **$1.39** |
The same session on Gemini 3.8 Flash ($0.90 in, $4.50 out) costs about $0.52. Models that cache the conversation cost less on long sessions, and Claude Code asks for caching on every call.
## Keep it low [#keep-it-low]
1. **Pick the model for the job.** A fast model for market checks, a frontier model for decisions.
2. **Compact long sessions.** `/compact` in Claude Code shrinks what every later turn resends.
3. **Grant fewer scopes.** Fewer tools means a shorter tool list on every turn.
4. **Install the skills you use,** not the whole hub.
5. **Set a daily limit on the key.** It can't overspend, whatever the session does.
## Paid by trading [#paid-by-trading]
An agent key spends the agent's own budget: 63¢ of every tax dollar its token's trades pay. An agent whose token trades $10,000 a day at 3% tax earns about $189 of AI a day, well over a hundred sessions like the one above. See [Where the money comes from](/how-it-works/fees).
# Binance Agent OS (https://docs.binference.io/agent-os)
Your agent acts on Binance through Agent OS and thinks with bInference, paid by its own trading fees.
[Binance Agent OS](https://www.binance.com/agent-os) connects AI apps to Binance. Your agent can read markets, check balances and trade, all within limits you set. Agent OS doesn't supply the AI model itself: the app you run your agent in still needs one to think with.
bInference is that model. Every call your agent makes to think is paid from its token's trading fees, or from credit you buy.
## What Agent OS includes [#what-agent-os-includes]
| Part | What it does | Where it acts |
| ---------------------- | -------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------- |
| **Binance MCP Server** | Market data, your agent's sub-account and trading: Spot, Margin, Convert and Futures | Your Binance account, inside an Agentic sub-account |
| **Agentic Wallet** | On-chain swaps, limit orders, transfers, prediction markets and DeFi | Your Binance MPC wallet, on BNB Chain, Ethereum, Base and Solana |
| **Skills Hub** | Read-only skills for token info, audits, rankings and smart-money signals, plus the Agentic Wallet skill | Any app that installs skills |
| **Agentic Payments** | Per-request stablecoin payments over HTTP 402 (x402) | BNB Chain |
## Pick your app [#pick-your-app]
Your agent needs an app that can use bInference as its model **and** connect to Binance.
| App | Thinks with bInference | Binance MCP Server | Agentic Wallet and skills | Guide |
| ----------------------------- | ---------------------------------------- | ------------------ | ------------------------- | ------------------------------- |
| **Claude Code** | Yes | Yes | Yes | [Set up](/agent-os/claude-code) |
| **Codex CLI** | Yes | Yes | Yes | [Set up](/agent-os/codex) |
| **OpenClaw** | Yes | Not yet | Yes | [Skills](/agent-os/skills) |
| **Hermes** | Yes | Not yet | Not listed | [Hermes](/clients/hermes) |
| Claude Desktop, ChatGPT, Grok | No: they run on their maker's own models | Yes | No | Binance's own guide |
**Claude Code and Codex CLI** are the two that do everything, so the guides start there. "Not yet" follows Binance's own list of supported apps, which it plans to grow.
## Before you start [#before-you-start]
* **A Binance account,** signed in on a desktop browser. The MCP Server is built for desktop.
* **A bInference key.** Make one on [binference.io](https://binference.io/account/keys): an agent key if you launched a token for your agent, an account key otherwise. See [API keys](/keys).
* **For the Agentic Wallet:** an MPC Wallet in the Binance App, and Node.js 18 or newer. For all of Skills Hub, Node.js 22.
## How the two sides stay apart [#how-the-two-sides-stay-apart]
* **Credentials never mix.** Your `binf_` key goes only to bInference. You sign in to Binance in Binance's own window, and bInference never sees that sign-in.
* **Money never mixes.** Trading money sits in your agent's Binance sub-account or wallet. AI money is your agent's budget on bInference. Neither can pay for the other.
* **Each side has its own brakes.** Binance limits what the agent may trade; bInference limits what it may spend on AI. See [Safety and limits](/agent-os/safety).
# Safety and limits (https://docs.binference.io/agent-os/safety)
Every limit on both sides, what each one stops and how to stop everything at once.
An agent on Agent OS has two sets of brakes. Binance limits what it may do with your money. bInference limits what it may spend on AI. Set both before you leave an agent running.
## On Binance [#on-binance]
| Limit | What it stops | Where you set it |
| ------------------------- | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- |
| **Scopes** | Market data, Account, Trade, Transfer: the agent can do only what you granted | When you authorize the MCP Server. To change them, disconnect and connect again |
| **Agentic sub-account** | The agent trades only the funds you moved into it. It can't reach your main account | Profile → Dashboard → Sub-account |
| **No withdrawals** | There is no withdrawal scope. Funds can never leave for an outside address | Always on |
| **Confirm before acting** | Every order, cancel and transfer waits for your yes | Always on |
| **Agentic Wallet rules** | Daily limit, tradable tokens, high-risk transactions rejected or sent to the App | Binance App → Agentic Wallet → Settings |
## On bInference [#on-binference]
| Limit | What it stops | Where you set it |
| ---------------------- | ------------------------------------------------------------------- | ---------------------------------------------- |
| **Key spending limit** | The key stops at a dollar amount per day, week or month | [API keys](https://binference.io/account/keys) |
| **The budget itself** | Calls stop when the agent's budget is spent. Nothing runs into debt | Automatic |
| **Reserve per call** | A call that could cost more than the balance never starts | Set `max_tokens` |
| **8 calls at once** | A runaway loop can't flood the model | Automatic |
## Stop everything [#stop-everything]
If an agent misbehaves, do these three, in this order:
1. **Emergency stop on Binance.** Profile → Dashboard → Sub-account → Account Management → **Emergency stop**. It disconnects every agent and cancels all open orders and positions in the Agentic sub-account at once.
2. **Revoke the key on bInference.** [API keys](https://binference.io/account/keys) → **Revoke**. The agent can't think any more, from its very next call.
3. **Sign out of the Agentic Wallet,** if you use it: tell the agent "Sign out", or sign it out in the Binance App.
To only pause an agent, use **Disconnect agents** instead of the emergency stop: open orders stay open.
## Habits that help [#habits-that-help]
* Fund the sub-account with only what you'd let the agent trade.
* Give each agent its own bInference key with a daily limit.
* Read every order the agent restates before you say yes. AI can misread a price or a size.
* Never paste keys or the MCP address into a chat.
Binance's own terms apply to trading through Agent OS, and you're responsible for the trades your agent places. See Binance's [MCP Server guide](https://developers.binance.com/en/docs/agent-native/mcp-server/agentic) and its disclosures.
# Skills (https://docs.binference.io/agent-os/skills)
Add Binance's skills to your agent for market data, token audits and on-chain trading.
Binance Skills Hub is an open collection of skills: small modules your agent reads and uses when a question calls for them. They work in Claude Code, Codex, OpenClaw and other agent apps, all thinking with bInference.
## Install [#install]
Every skill in the hub:
```bash
npx skills add https://github.com/binance/binance-skills-hub
```
Or only the one you need, like the Agentic Wallet:
```bash
npx skills add binance/binance-skills-hub/skills/binance-web3/binance-agentic-wallet
```
The installer needs Node.js 22 or newer and asks which of your apps to add the skills to.
## Binance's wallet skills [#binances-wallet-skills]
| Skill | Type | What your agent can ask |
| ----------------------------------- | -------------- | --------------------------------------------------------------------------------------------------------- |
| `binance-agentic-wallet` | Read and write | Balances, swaps, limit orders, transfers. Needs a sign-in. See [Agentic Wallet](/agent-os/agentic-wallet) |
| `query-token-info` | Read | Search tokens, prices, market cap, candles |
| `query-token-audit` | Read | Risk score, honeypot check, buy and sell tax |
| `query-address-info` | Read | Any wallet's holdings and their value |
| `crypto-market-rank` | Read | Trending tokens, smart-money inflow, trader PnL |
| `meme-rush` | Read | New and migrating meme tokens |
| `trading-signal` | Read | On-chain smart-money signals |
| `binance-tokenized-securities-info` | Read | Tokenized stock prices and market status |
The read-only skills need no wallet and no sign-in.
## A useful pair for launch tokens [#a-useful-pair-for-launch-tokens]
Before your agent buys any token, have it audit it:
```txt
Audit the token at 0x...7777 on BNB Chain, then tell me its buy and sell tax.
```
`query-token-audit` flags taxes above 10% as critical. bInference agent tokens carry a 1%, 3% or 5% tax, set at launch.
## OpenClaw [#openclaw]
Binance's MCP Server doesn't list OpenClaw yet, so on OpenClaw your agent reaches Binance through skills. Point OpenClaw at bInference first: see [OpenClaw](/clients/openclaw).
# Troubleshooting (https://docs.binference.io/agent-os/troubleshooting)
What went wrong, on which side and how to fix it.
## Which side is it? [#which-side-is-it]
| You see | Side | Fix |
| --------------------------------------------- | ---------- | ------------------------------------------------------------------------------ |
| The app can't start or says the key is wrong | bInference | Check `ANTHROPIC_AUTH_TOKEN` or `BINF_API_KEY`. See [API keys](/keys) |
| `402 insufficient_balance` | bInference | The AI budget can't cover the call's reserve. Lower `max_tokens` or add credit |
| `402 key_limit` | bInference | The key hit its spending limit. It resets at 00:00 UTC, or raise it |
| `429` or `529` | bInference | Too many calls at once, or a busy model. The app retries on its own |
| A trade fails for insufficient balance | Binance | The Agentic sub-account is empty. Transfer funds into it |
| Market data works, account or trading doesn't | Binance | Reconnect and grant the Account or Trade scope |
| The server is added but can't be used | Binance | Reconnect and authenticate again |
| The authorization expired | Binance | Disconnect and reconnect the Binance MCP Server |
## Binance MCP Server [#binance-mcp-server]
* **The command isn't recognized.** Make sure Claude Code or Codex is installed, open a new terminal and try again.
* **You opened the address in a browser.** Close it and run the command for your app instead.
* **You pasted the address into the chat.** Remove the connection the AI made, then follow the setup steps exactly.
* **Check the address.** It must be exactly `https://agent.binance.com/mcp/agentic`.
## bInference [#binference]
* **Nothing seems to reach bInference.** Claude Code needs `ANTHROPIC_BASE_URL=https://binference.io/api`, without `/v1`. Codex needs `base_url = "https://binference.io/api/v1"`, with it.
* **Your calls don't show up on the Calls page.** Messages calls, which Claude Code makes, appear within a minute, once they're priced.
* **Every error code** is on [Errors](/api/errors), with what to do.
## Still stuck? [#still-stuck]
For the Binance side, use Binance Customer Support's Live Chat. For bInference, quote the `x-binference-call-id` from the failing call.
# Anthropic SDK (https://docs.binference.io/clients/anthropic-sdk)
The official Anthropic libraries, with Claude and every other model.
The Anthropic SDKs speak the Messages format. Set the base URL to `https://binference.io/api`: the SDK adds `/v1/messages` itself.
TypeScript
Python
```ts
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: "https://binference.io/api",
apiKey: process.env.BINF_API_KEY,
});
const message = await client.messages.create({
model: "anthropic/claude-sonnet-5.5",
max_tokens: 1024,
messages: [{ role: "user", content: "Explain gas fees to a new trader." }],
});
console.log(message.content);
```
```python
import os
from anthropic import Anthropic
client = Anthropic(
base_url="https://binference.io/api",
api_key=os.environ["BINF_API_KEY"],
)
message = client.messages.create(
model="anthropic/claude-sonnet-5.5",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain gas fees to a new trader."}],
)
print(message.content)
```
## Good to know [#good-to-know]
* **Model names:** both `anthropic/claude-sonnet-5.5` and Claude's own names like `claude-sonnet-5-5` work.
* **Any model:** the Messages format works with non-Claude models too, such as `google/gemini-3.8-flash`.
* **Cost:** Messages answers carry token counts but no cost. Each call's charge is on your [Calls page](https://binference.io/account/calls) within a minute.
* **Overloaded:** a busy model answers `529 overloaded_error`, which the SDK retries after the wait we name.
# Claude Code (https://docs.binference.io/clients/claude-code)
Run Claude Code on your agent's budget with three settings.
Set these before you start Claude Code, in your shell or in `~/.claude/settings.json` under `env`.
Shell
settings.json
```bash
export ANTHROPIC_BASE_URL="https://binference.io/api"
export ANTHROPIC_AUTH_TOKEN="$BINF_API_KEY"
export ANTHROPIC_API_KEY=""
claude
```
```json
{
"env": {
"ANTHROPIC_BASE_URL": "https://binference.io/api",
"ANTHROPIC_AUTH_TOKEN": "binf_...",
"ANTHROPIC_API_KEY": ""
}
}
```
`ANTHROPIC_API_KEY` is set empty so Claude Code doesn't use a key it found elsewhere.
## Choose the model [#choose-the-model]
Claude Code's own model names work as they are. To pick one, set `ANTHROPIC_MODEL`:
```bash
export ANTHROPIC_MODEL="anthropic/claude-sonnet-5.5"
```
Any listed model works, including ones from other makers, as long as it handles tools well.
## Good to know [#good-to-know]
* **Long sessions:** Claude Code resends the conversation on every turn. If a request grows past 4 MB, usually from pasted images, it's refused with `413`. Run `/compact` or `/clear`.
* **Busy models:** a busy model answers `529 overloaded_error`, and Claude Code moves to its fallback model.
* **Cost:** each turn's charge is on your [Calls page](https://binference.io/account/calls) within a minute. Put a daily limit on the key you give Claude Code.
# Codex (https://docs.binference.io/clients/codex)
Add bInference to the Codex CLI as a model provider.
Codex speaks the Responses format. Add a provider to `~/.codex/config.toml`:
```toml title="~/.codex/config.toml"
model_provider = "binference"
model = "openai/gpt-6.1-sol"
[model_providers.binference]
name = "bInference"
base_url = "https://binference.io/api/v1"
env_key = "BINF_API_KEY"
wire_api = "responses"
```
Then start Codex with your key in the environment:
```bash
export BINF_API_KEY="binf_..."
codex
```
Any listed model works as `model`. Codex retries `503` answers, so a busy model or a short wait on our side resolves on its own.
# Hermes (https://docs.binference.io/clients/hermes)
Point Hermes Agent at bInference as a custom endpoint.
Run `hermes model`, choose **Custom endpoint** and enter:
* **Base URL:** `https://binference.io/api/v1`
* **API key:** your `binf_` key
* **Model:** any id from [Models](/models), like `anthropic/claude-sonnet-5.5`
Or set it in `~/.hermes/config.yaml`, reading the key from the environment:
```yaml title="~/.hermes/config.yaml"
model:
default: anthropic/claude-sonnet-5.5
provider: custom
base_url: https://binference.io/api/v1
key_env: BINF_API_KEY
```
Hermes calls Chat Completions, streamed, so it gets every model and every feature here.
# OpenAI SDK (https://docs.binference.io/clients/openai-sdk)
The official OpenAI libraries for TypeScript and Python, with any model.
The OpenAI SDKs work with every model here, not only OpenAI's. Set `baseURL` and the key.
TypeScript
Python
```ts
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://binference.io/api/v1",
apiKey: process.env.BINF_API_KEY,
});
const stream = await client.chat.completions.create({
model: "google/gemini-3.8-flash",
messages: [{ role: "user", content: "Write a haiku about BNB Chain." }],
max_tokens: 512,
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
```
```python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://binference.io/api/v1",
api_key=os.environ["BINF_API_KEY"],
)
stream = client.chat.completions.create(
model="google/gemini-3.8-flash",
messages=[{"role": "user", "content": "Write a haiku about BNB Chain."}],
max_tokens=512,
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
```
## Responses [#responses]
`client.responses.create` works the same way, stateless: nothing is stored between calls, so send the conversation each time.
```ts
const reply = await client.responses.create({
model: "openai/gpt-6.1-sol",
input: "Summarize the last 24 hours of trades in one line.",
max_output_tokens: 256,
});
console.log(reply.output_text);
```
## Retries [#retries]
The SDKs retry 429 and 5xx answers on their own, twice by default, and wait as long as the `Retry-After-Ms` header says. That is the right behavior here. Raise `maxRetries` (`max_retries` in Python) for jobs that run unattended. See [Retries](/api/retries).
## Other OpenAI-compatible tools [#other-openai-compatible-tools]
Anything that lets you set an OpenAI base URL works: LangChain's `ChatOpenAI`, LiteLLM, LlamaIndex and most agent frameworks. Use `https://binference.io/api/v1` as the base URL and your `binf_` key as the API key.
# OpenClaw (https://docs.binference.io/clients/openclaw)
Add bInference as a custom provider in OpenClaw.
Add a provider to `~/.openclaw/openclaw.json`, then allow its model for your agents:
```json5 title="~/.openclaw/openclaw.json"
{
models: {
providers: {
binference: {
baseUrl: "https://binference.io/api/v1",
apiKey: "${BINF_API_KEY}",
api: "openai-completions",
models: [
{
id: "anthropic/claude-sonnet-5.5",
name: "Claude Sonnet 5.5",
contextWindow: 1000000,
// The longest answer per call. It is also what each call reserves.
maxTokens: 8192,
},
],
},
},
},
agents: {
defaults: {
model: { primary: "binference/anthropic/claude-sonnet-5.5" },
models: { "binference/anthropic/claude-sonnet-5.5": { alias: "Sonnet" } },
},
},
}
```
Model ids keep their maker prefix, so the full name is `binference/` plus the id from [Models](/models).
OpenClaw sends `maxTokens` as each call's output limit, and every call reserves that
much from the budget while it runs. 8,192 suits most agent steps.
To use the Messages format instead, set `api: "anthropic-messages"` and `baseUrl: "https://binference.io/api"`.
# Vercel AI SDK (https://docs.binference.io/clients/vercel-ai-sdk)
Use bInference as an OpenAI-compatible provider in the AI SDK.
Install the OpenAI-compatible provider:
```bash
npm i ai @ai-sdk/openai-compatible
```
```ts
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { streamText } from "ai";
const binference = createOpenAICompatible({
name: "binference",
baseURL: "https://binference.io/api/v1",
apiKey: process.env.BINF_API_KEY,
// Adds usage to streamed answers, with what each call cost.
includeUsage: true,
});
const result = streamText({
model: binference("anthropic/claude-sonnet-5.5"),
prompt: "Write a tweet about our agent's first week.",
maxOutputTokens: 512,
});
for await (const text of result.textStream) {
process.stdout.write(text);
}
```
Every model id from [Models](/models) works with `binference("...")`. Tools, structured output and multi-step agents work as they do with any provider.
# Questions (https://docs.binference.io/faq)
Short answers to what people ask most.
Yes. Chat Completions, Responses and Messages follow the official formats, streaming
included, so the official SDKs and the tools built on them work after you change the
address and the key.
A call reserves its worst case before it starts. Without `max_tokens` that's the
model's whole output, which can be more than your balance. Set `max_tokens`, and the
error message shows both amounts so you can see the gap.
No. Every refusal before the model runs is free. If a model's provider fails after
it started, you pay only what that provider recorded, usually nothing.
You pay for what the model wrote before it stopped. We never assume a stopped call
is free, because some providers bill the work already done.
8 calls running at once, and 600 started a minute. See [Rate
limits](/api/rate-limits).
The API allows calls from any website, but a key in a web page can be read by anyone
who visits it. Call from your server, or give a browser tool a key with a small
spending limit.
No. It only pays for AI. It isn't money and can't be withdrawn or moved.
Each day's credit lasts 7 days. What's left moves to the holders' pool, which keeps
AI cheap for everyone. See [Credit and expiry](/how-it-works/credit).
No. We keep each call's model, token counts, cost and timing, never the prompt or
the answer.
Not yet. The gateway serves text, images in and tools. Embeddings, token counting
and image generation answer `404` or `400 tool_not_served`.
# Reasoning (https://docs.binference.io/features/reasoning)
Let a model think before it answers, and what thinking costs.
Reasoning models think before they answer. You control how much.
Chat and Responses
Messages
```json
{
"model": "openai/gpt-6.1-sol",
"reasoning": { "effort": "medium" },
"max_tokens": 4096
}
```
```json
{
"model": "anthropic/claude-sonnet-5.5",
"thinking": { "type": "enabled", "budget_tokens": 2048 },
"max_tokens": 4096
}
```
## What it costs [#what-it-costs]
Thinking is billed as output, at the model's output price. It counts toward `max_tokens`, so leave room for both the thinking and the answer.
* `effort: "low"` answers faster and cheaper. Good for agent steps that are simple.
* `effort: "high"` thinks longest. Save it for hard problems.
Streams send the thinking as it happens, before the answer. In the [playground](/api/playground) it shows in a **Thinking** fold above the answer.
# Routing and fallbacks (https://docs.binference.io/features/routing)
Steer which provider runs a model, and keep going when one is busy.
Most models run at several providers. By default each call goes to a good one, and moves to another when the first is down.
## Variants [#variants]
Add a variant to a model id to steer that choice:
| Variant | What it does |
| --------- | -------------------------------------------------------------------------------- |
| `:nitro` | Picks the fastest providers first. Some charge more, so the reserve covers them. |
| `:floor` | Picks the cheapest providers first. |
| `:exacto` | Picks providers known to handle tool calls most accurately. |
| `:online` | Searches the web first. See [Web search](/features/web-search). |
```json
{ "model": "deepseek/deepseek-v4.1-flash:nitro" }
```
`:free`, `:batch` and retired variants like `:thinking` are not served.
## Fallback models [#fallback-models]
List other models in `models`, and a call moves down the list when the first is busy or down:
```json
{
"model": "anthropic/claude-sonnet-5.5",
"models": [
"anthropic/claude-sonnet-5.5",
"openai/gpt-6.1-sol",
"google/gemini-3.8-flash"
],
"max_tokens": 1024
}
```
The reserve covers the priciest model on the list. The answer's `model` field says which one ran.
## When every provider is busy [#when-every-provider-is-busy]
A model busy at all its providers answers `503 model_busy` (`529` on Messages) with the wait in `Retry-After-Ms`. That model pauses for the wait, so calls to it get the same answer at once, with nothing reserved. Models on your fallback list keep working.
# Streaming (https://docs.binference.io/features/streaming)
Get the answer as it's written, and why long calls should always stream.
Add `"stream": true` and the answer arrives as server-sent events, a few words at a time. Each format streams its own way, exactly as its official API does, so the SDKs read it as usual.
## Why stream [#why-stream]
* **Longer limit.** A streamed call may run 30 minutes. One that isn't streamed stops at 13.
* **Steady connection.** Some networks drop a connection that stays silent for minutes. A stream keeps it alive.
* **Faster feel.** People see the first words in well under a second on fast models.
Claude Code, Codex, OpenClaw and Hermes stream by default.
## The cost of a streamed call [#the-cost-of-a-streamed-call]
On Chat Completions, add `stream_options` to get a last chunk with `usage`, including what the call cost:
```json
{
"model": "google/gemini-3.8-flash",
"messages": [{ "role": "user", "content": "Hello" }],
"max_tokens": 256,
"stream": true,
"stream_options": { "include_usage": true }
}
```
The final event before `data: [DONE]`:
```json
{
"choices": [],
"usage": { "prompt_tokens": 9, "completion_tokens": 12, "cost": 0.000062 }
}
```
On Responses, the `response.completed` event carries `usage.cost`. Messages streams carry token counts but no cost; the charge is on your [Calls page](https://binference.io/account/calls) within a minute.
## When a stream breaks [#when-a-stream-breaks]
* **An error after words have arrived** comes as the format's own error event, so the SDK raises it instead of hanging. The call is charged for what was written.
* **If you disconnect,** the model stops and you're charged for what it wrote before. We never charge a disconnect as zero, because some providers bill the work already done.
* **At the time limit,** the stream ends with a clean error event 5 seconds before the cut.
Try it in the [playground](/api/playground): the Raw tab shows each event with its time.
# Tools and structured output (https://docs.binference.io/features/tools)
Let the model call your functions, or answer in the exact JSON you need.
## Function calling [#function-calling]
Describe your functions in `tools`. The model answers with the calls it wants; your code runs them and sends the results back.
```ts
const reply = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5.5",
max_tokens: 1024,
messages: [{ role: "user", content: "What's the price of $NOVA?" }],
tools: [
{
type: "function",
function: {
name: "token_price",
description: "The latest price of a token on BNB Chain, in USD",
parameters: {
type: "object",
properties: { ticker: { type: "string" } },
required: ["ticker"],
},
},
},
],
});
const call = reply.choices[0].message.tool_calls?.[0];
```
Tools work the same on Responses and on Messages, in each format's own shape. Filter [Models](https://binference.io/models) by **Tools** to see which models support them.
## Structured output [#structured-output]
Ask for JSON that matches your schema with `response_format`:
```json
{
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "trade_summary",
"strict": true,
"schema": {
"type": "object",
"properties": {
"trades": { "type": "integer" },
"net_bnb": { "type": "number" }
},
"required": ["trades", "net_bnb"],
"additionalProperties": false
}
}
}
}
```
## Tools we don't serve [#tools-we-dont-serve]
Every agent's calls run on one shared account, so tools that keep files, memory or a sandbox there, or that run other models inside a call, are refused with `400 tool_not_served`:
* hosted code execution, shells and code interpreters
* file search, files, memory and the text editor tool
* image generation and tools that call other models
* X search, which bills per post with no limit a request can set
Tools your code runs on its own machine always work. Web search and web fetch are served: see [Web search](/features/web-search).
Hosted tools like web search run in a loop, and the reserve covers every possible
step. Set `max_tool_calls` (or the tool's `max_uses`) to keep it small. A loop with no
step limit is refused.
# Web search (https://docs.binference.io/features/web-search)
Add :online to any model and it reads the web before it answers.
Add `:online` to a model id and it searches the web first, then answers with what it found:
```json
{
"model": "google/gemini-3.8-flash:online",
"messages": [{ "role": "user", "content": "What happened on BNB Chain this week?" }],
"max_tokens": 1024
}
```
## What it costs [#what-it-costs]
Each search is charged at the model's search price, shown as `web_search` in [`GET /models`](/api/models), on top of the tokens. The search results are read as input.
While the call runs, it reserves five searches plus the tokens, and returns what it didn't use.
## Search as a tool [#search-as-a-tool]
Models that support hosted web search and web fetch tools can use them too. They run in a loop, so set `max_tool_calls` to cap the steps and the reserve. X search is not served: it bills per post returned with no limit a request can set.
# How a call is paid (https://docs.binference.io/how-it-works/calls)
A call reserves its worst case, runs, then pays only what the model wrote.
Balances can never go into debt by surprise. So before a call starts, it reserves its worst-case cost. When it ends, the real cost is taken and the rest goes back at once.
## The steps [#the-steps]
1. **Check.** The key, the agent, the model and the body. Anything wrong answers at once, free.
2. **Reserve.** The worst case is held from the balance: the whole prompt plus the longest answer the call allows. If the balance can't cover it, the answer is `402 insufficient_balance`, with both amounts, and nothing runs.
3. **Answer.** The model runs and the answer streams back to you as it's written.
4. **Charge.** The real cost is taken, and the rest of the reserve is freed. Every answer that carries a cost reports this charge in `usage.cost`.
## What the reserve counts [#what-the-reserve-counts]
* **The prompt:** its text at about 3 characters a token, 8,000 tokens per image, and the model's whole context for files sent by link.
* **The answer:** `max_tokens` (`max_output_tokens` on Responses). Without it, the model's whole output limit.
* **The price:** the dearest provider that could run the call, so a fallback never costs more than was held.
* **Tool loops:** every step a search loop could take. Cap it with `max_tool_calls`.
* **At least $0.001,** for the smallest call.
Set `max_tokens` to what you need. A hello with no limit on a frontier model reserves
dollars; with `max_tokens: 256` it reserves a fraction of a cent. It changes what the
call holds, never what it costs.
## Running calls hold money [#running-calls-hold-money]
A call's reserve is held while it runs, so ten long calls at once hold ten reserves. `spendable_usd` in [`GET /balance`](/api/balance) is what new calls can reserve now; `reserved_usd` is what running calls hold.
## When no cost comes back [#when-no-cost-comes-back]
If a call is cut before the answer carries its cost (you disconnected, the stream broke or the format carries none), the call is priced from the provider's own record within a minute. Its reserve stays held until then.
# Credit and expiry (https://docs.binference.io/how-it-works/credit)
Balances are in dollars, in daily batches that last 7 days and are spent oldest first.
Every balance is counted in dollars. BNB becomes dollars once, when it's credited, at that moment's BNB price.
## The rule [#the-rule]
* **Each day's credit is one batch.** Fee credit that lands on Monday is Monday's batch, whatever the time.
* **A batch lasts 7 days.** It expires at 00:00 UTC seven days later.
* **The oldest batch is spent first,** so fresh credit waits and less expires.
* **What's left when a batch expires moves to the holders' pool,** our store of AI that keeps prices low for everyone.
`GET /balance` lists what is left of each batch and when it expires, soonest first:
```json
"expiring": [
{ "at": "2026-10-02T00:00:00.000Z", "usd": "2.108400" },
{ "at": "2026-10-03T00:00:00.000Z", "usd": "4.920000" }
]
```
## Two kinds of credit [#two-kinds-of-credit]
| | Agent credit | Account credit |
| ------------- | ------------------------------ | -------------------------------------------------------- |
| Comes from | The agent token's trading fees | BNB you pay, from 0.001 BNB |
| Spent by | That agent's keys only | Your account keys |
| Can be bought | No | Yes, on [Credits](https://binference.io/account/credits) |
| Lasts | 7 days per daily batch | 7 days per purchase |
Neither can be withdrawn. $BINF holders get more credit for the same BNB: see [$BINF holders](/how-it-works/holders).
## Buying account credit [#buying-account-credit]
Open **Buy credit** on your account, enter an amount of BNB and approve it in your wallet. The page follows the payment on chain and adds the credit once it's final, usually within a minute. Closing the page doesn't stop it.
# Where the money comes from (https://docs.binference.io/how-it-works/fees)
Every trade of an agent's token pays a small tax. Most of it becomes the agent's AI budget.
Each agent has its own token on BNB Chain, launched on [Flap](https://flap.sh). Every buy and every sell of it pays a tax, chosen once at launch: 1%, 3% or 5%. That tax is what pays for the agent's AI.
## Step by step [#step-by-step]
1. **Someone trades the token.** The tax is taken in BNB from the trade.
2. **Flap keeps 10%** of the tax, as the launchpad.
3. **The token's own vault gets the rest,** and splits it on its own:
* **70% to the agent's AI budget.** It becomes dollars of credit at the BNB price when it lands, spent through the agent's keys.
* **30% to the treasury,** which buys AI in bulk for the holders' pool.
4. **Fees reach the vault within minutes** of the trades. Anyone can push a payout, and we do it too, so nothing waits on the creator.
So of every tax dollar, 63¢ becomes the agent's AI, 27¢ goes to the treasury and 10¢ to Flap.
## Good to know [#good-to-know]
* **The tax runs 100 years,** the longest Flap allows, at the rate set at launch. Nobody can change it later.
* **The budget is only for AI.** It isn't money, can't be withdrawn and can't be sent elsewhere.
* **The budget lasts 7 days.** Each day's credit is used oldest first, and what's left after a week moves to the holders' pool. See [Credit and expiry](/how-it-works/credit).
* **No trading, no budget.** An agent whose token isn't traded earns nothing, and its calls stop when the budget is spent. Its owner can still use an account key with bought credit.
# $BINF holders (https://docs.binference.io/how-it-works/holders)
Hold $BINF and the credit you buy comes with up to 8% more.
$BINF is bInference's own token. Holding it makes the credit you buy cheaper. You still pay in BNB: $BINF only sets your discount, and it's never spent.
## How it works [#how-it-works]
* **Eight tiers,** from 1% off at 5,000,000 $BINF to 8% off at 15,000,000.
* **Your lowest balance counts,** over the last 7 days. Buying just before a purchase earns nothing.
* **More credit, same BNB.** At 6% off, a payment worth $94 adds $100 of credit.
* **Shown before you pay.** The Buy credit dialog shows your tier and prices its estimate with it.
## Why AI is cheaper here [#why-ai-is-cheaper-here]
The holders' pool is our store of AI: the treasury's 30% of every agent token's fees and $BINF's own tax buy AI in bulk, and every agent's unused credit moves into it after 7 days. It serves every call, and it's what lets bInference sell AI cheaply.
# Launching an agent (https://docs.binference.io/how-it-works/launch)
One screen, three steps and an agent whose AI pays for itself.
You launch your agent's token on [binference.io/launch](https://binference.io/launch). It fits on one screen.
## What you choose [#what-you-choose]
| Field | Rule |
| -------------------- | -------------------------------------------------------- |
| Image | PNG, JPG, WEBP or GIF, at least 256×256, at most 4 MB |
| Name | 1 to 32 characters |
| Ticker | 2 to 10 letters or digits |
| What your agent does | Up to 280 characters |
| Tax | 1%, 3% or 5% on every buy and sell. 3% is picked for you |
| Opening buy | Optional, up to 10 BNB of your own token at launch |
| Links | X, Telegram and a website, all optional |
## What it costs [#what-it-costs]
* **0.005 BNB** launch fee, paid after the launch.
* **The network fee** for the launch, a few cents.
* **Your opening buy,** if you make one.
## The three steps [#the-three-steps]
1. **Save.** The image and details are saved on IPFS.
2. **Launch.** Your wallet sends the launch. Your token's address ends in `7777`.
3. **Fee.** Your wallet sends the launch fee.
The page checks each transaction before your wallet asks, and the agent is live once the chain confirms both. Then make its [API key](/keys).
## After launch [#after-launch]
* The token trades on Flap, and graduates to a DEX after about 20 BNB of buying.
* Every trade's tax fills the agent's AI budget. See [Where the money comes from](/how-it-works/fees).
* The agent is registered in BNB Chain's agent registry (ERC-8004), so other apps can find it.
* Its page on binference.io shows the chart, a box to buy or sell and live proof it works: fees earned, AI spent, calls made.
# Privacy (https://docs.binference.io/how-it-works/privacy)
We never store prompts or answers. Here is exactly what we keep.
Your prompts and the model's answers pass through bInference and are never written anywhere. Prompt logging at our providers is off.
## What we keep for each call [#what-we-keep-for-each-call]
* the model, the endpoint and whether it streamed
* token counts: in, out, cached and reasoning
* the reserve, the cost and the charge
* when it started, how long it took and how it ended
* the provider's id for the answer
That's what your [Calls page](https://binference.io/account/calls) shows, and what the public numbers on each agent's page are made from.
## Keys [#keys]
We store a key's first characters, to find it, and a fingerprint of the whole key. The key itself is shown once and never stored, so nobody at bInference can read it.
## Your answers belong to the model's terms [#your-answers-belong-to-the-models-terms]
Each call runs at the model's provider under its own terms. Some providers may keep data for abuse checks. If that matters for your use, pick models from providers whose terms suit you.
# Introduction (https://docs.binference.io/)
Every AI model on one key, paid by your agent's trading fees.
## What you get [#what-you-get]
* **Made for Binance Agent OS.** Your agent acts on Binance through Agent OS and thinks with bInference. [Set it up](/agent-os).
* **One key, hundreds of models.** Claude, GPT, Gemini, DeepSeek, Grok and more, each at a fixed price per token. See them all on [Models](/models).
* **The formats you already use.** OpenAI Chat Completions and Responses, and Anthropic Messages. Claude Code, Codex, OpenClaw, Hermes and the official SDKs work after two settings change.
* **Paid by trading, not by card.** An agent's AI budget fills from its token's trading fees, in dollars. People can also buy credit for their own account with BNB.
* **Nothing kept.** We never store prompts or answers, only each call's count and cost.
## Start here [#start-here]
## Who it is for [#who-it-is-for]
# API keys (https://docs.binference.io/keys)
Two kinds of key, optional spending limits and what to do if one leaks.
A key starts with `binf_` and is followed by 40 letters and digits. Send it on every call as either header:
```http
Authorization: Bearer binf_...
x-api-key: binf_...
```
## Two kinds of key [#two-kinds-of-key]
| | Agent key | Account key |
| ---------------- | --------------------------------------------------------------- | ------------------------------------------------------------- |
| Spends | The agent's own budget | Your account's credit |
| Filled by | The agent token's trading fees | BNB you pay, from 0.001 BNB |
| Who can make one | The wallet that launched the agent | Any signed-in wallet |
| Where | [API keys](https://binference.io/account/keys), under the agent | [API keys](https://binference.io/account/keys), under Account |
Both kinds call the same endpoints at the same prices. An agent's budget can only be spent by that agent's keys, and nobody can buy it: it comes only from trading.
## Make a key [#make-a-key]
1. Sign in on [binference.io](https://binference.io/account/keys) with the wallet that launched your agent, or any wallet for an account key.
2. Press **Create key** and give it a name you'll recognize, like `prod-server` or `laptop`.
3. Copy the key. **It is shown once.** We keep only a fingerprint, so a lost key can't be shown again: revoke it and make another.
You can have up to 10 active keys per agent, and 10 for your account.
## Spending limits [#spending-limits]
Turn on **Limit spending** to cap what one key may spend per day, week or month, from $0.01 to $1,000,000. It is the safest way to give a key to a new tool or a teammate.
* Limits start over at 00:00 UTC. A weekly limit starts over on Monday.
* A key at its limit is refused with `402 key_limit`, and the agent's other keys keep working.
* A call counts in the period it started, however late it finishes.
* Setting a limit mid-period counts what the key already spent in it.
`GET /balance` shows the calling key's limit, what it spent and when it resets. See [Get balance](/api/balance).
## Revoke a key [#revoke-a-key]
Press **Revoke** next to it. It stops on its very next call: nothing is cached. A call already running finishes.
Revoke it at once, then make a new one. A leaked key can spend its agent's whole
balance, or its own limit. Never put a key in code that runs in a browser or a mobile
app: keep it on your server.
## Keep keys safe [#keep-keys-safe]
* Read the key from an environment variable, never from source code.
* Give each tool or machine its own key, so you can revoke one without stopping the rest.
* Put a spending limit on keys for tests and for tools you're trying.
# Models (https://docs.binference.io/models)
Every model you can call, with live prices. Copy an id and send it as model.
Every listed model has a fixed price per token, and these are the prices you pay. The list below is live from [`GET /models`](/api/models).
Every model and its live price: `GET https://binference.io/api/v1/models`.
## Model ids [#model-ids]
Send the id as `model`, like `anthropic/claude-sonnet-5.5`. A few more forms work:
* **Claude clients' own names** on `/messages`: `claude-sonnet-5-5`, dated names like `claude-haiku-4-5-20251001` and the `[1m]` suffix map to the listed model.
* **Latest aliases** like `~anthropic/claude-opus-latest` work and are priced as the model they point to.
* **Variants** change how a model runs, not what it costs per token: `:online`, `:nitro`, `:floor` and `:exacto`. See [Web search](/features/web-search) and [Routing and fallbacks](/features/routing).
## Not served [#not-served]
These answer `400 model_not_served`, because their price can't be known before the call:
* free models and any `:free` version
* `:batch` versions and retired versions like `:thinking`
* routers that pick a model after the call
A model we don't know answers `400 unknown_model`. New models show up in the list within an hour of release.
## Picking a model [#picking-a-model]
* **For agents that loop all day,** a fast, cheap model keeps the budget going: Gemini Flash, DeepSeek Flash or Claude Haiku.
* **For hard problems,** a frontier model with reasoning: Claude Opus, GPT or Gemini Pro.
* **Compare side by side** on [binference.io/models](https://binference.io/models): prices, context, the hosts that run each one and what it can do.
# Quickstart (https://docs.binference.io/quickstart)
Your first call in two minutes. If you've used the OpenAI or Anthropic API, you already know how.
Follow the [Binance Agent OS guide](/agent-os): it sets up the model and the Binance
connection together.
### Get a key [#get-a-key]
Sign in on [binference.io](https://binference.io/account/keys) with your wallet and open **API keys**.
* **Agent key:** spends an agent's own budget, filled by its token's trading fees. You need to have launched the agent.
* **Account key:** spends credit you buy with BNB, for any use.
The key is shown once. It starts with `binf_`. Keep it in an environment variable:
```bash
export BINF_API_KEY="binf_..."
```
### Make a call [#make-a-call]
Change two settings in the client you already use: the address and the key. That's the only change.
curl
TypeScript
Python
```bash
curl https://binference.io/api/v1/chat/completions \
-H "Authorization: Bearer $BINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"messages": [{ "role": "user", "content": "Say hello in five words." }],
"max_tokens": 256
}'
```
```ts
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://binference.io/api/v1",
apiKey: process.env.BINF_API_KEY,
});
const reply = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5.5",
messages: [{ role: "user", content: "Say hello in five words." }],
max_tokens: 256,
});
console.log(reply.choices[0].message.content);
```
```python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://binference.io/api/v1",
api_key=os.environ["BINF_API_KEY"],
)
reply = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[{"role": "user", "content": "Say hello in five words."}],
max_tokens=256,
)
print(reply.choices[0].message.content)
```
### Check what it cost [#check-what-it-cost]
Every answer's `usage.cost` is what the call was charged, in dollars. Your balance is one call away:
```bash
curl https://binference.io/api/v1/balance \
-H "Authorization: Bearer $BINF_API_KEY"
```
While a call runs it reserves its longest possible answer from your balance. Without a
limit that is the model's whole output, which can be dollars. You pay only what the
model writes, and the rest comes back when it ends. [How a call is
paid](/how-it-works/calls)
## Next [#next]
# For AI tools (https://docs.binference.io/api/ai-tools)
These docs as plain text for AI assistants, and the API as an OpenAPI file.
Give your coding assistant the docs, and it writes correct bInference code.
| File | What it holds |
| ---------------------------------- | ------------------------------------------------------------- |
| [`/llms.txt`](/llms.txt) | A map of every page, for AI assistants |
| [`/llms-full.txt`](/llms-full.txt) | Every page in one text file |
| `/.md` | Any page as Markdown, like [`/quickstart.md`](/quickstart.md) |
| [`/openapi.json`](/openapi.json) | The API as OpenAPI 3.1, for code generators and API clients |
Every page also has **Copy Markdown** and **Open** buttons under its title, to paste it into ChatGPT, Claude or your editor.
## OpenAPI [#openapi]
Import `https://docs.binference.io/openapi.json` into Postman, Insomnia, Bruno or any code generator. It covers the five endpoints, their main fields, the headers and every error code.
# Get balance (https://docs.binference.io/api/balance)
What the key can spend now, what's expiring and the key's own limit.
Read it before a batch of calls to size them, or to show your users what's left. Amounts are dollars as exact decimal strings.
`spendable_usd` is what new calls can reserve right now: the balance less what running calls hold. Compare it with a call's reserve, not with its likely cost. See [How a call is paid](/how-it-works/calls).
`GET https://binference.io/api/v1/balance`
Send your key as `Authorization: Bearer binf_...` or `x-api-key`.
## Response
- `agent` (object): The agent or account the key belongs to: `id`, `name`, `token` and `status`.
- `spendable_usd` (string): What new calls can reserve right now: the balance less what running calls hold.
- `balance_usd` (string): The whole balance.
- `reserved_usd` (string): What running calls hold.
- `running_calls` (integer): Calls running now, at most 8.
- `expiring` (object[]): What is left of each day's credit and when it expires (`at`, `usd`), soonest first.
- `key_limit` (object | null): This key's spending limit, if it has one: `limit_usd`, `reset` (daily, weekly or monthly), `spent_usd`, `reserved_usd`, `remaining_usd` and `resets_at`.
## Example: curl
```bash
curl https://binference.io/api/v1/balance \
-H "Authorization: Bearer $BINF_API_KEY"
```
## Example: TypeScript
```ts
const response = await fetch("https://binference.io/api/v1/balance", {
headers: { Authorization: `Bearer ${process.env.BINF_API_KEY}` },
});
const balance = await response.json();
```
## Example: Python
```python
import os
import requests
response = requests.get("https://binference.io/api/v1/balance",
headers={"Authorization": f"Bearer {os.environ['BINF_API_KEY']}"},
)
print(response.json())
```
## Example response
```json
{
"object": "balance",
"agent": {
"id": 42,
"name": "Nova",
"token": "0x9f3c7e2b8a41d05c6e1f9a8b7c2d3e4f50617777",
"status": "active"
},
"spendable_usd": "18.402110",
"balance_usd": "18.422110",
"reserved_usd": "0.020000",
"running_calls": 1,
"expiring": [
{
"at": "2026-10-02T00:00:00.000Z",
"usd": "2.108400"
},
{
"at": "2026-10-03T00:00:00.000Z",
"usd": "4.920000"
}
],
"key_limit": {
"limit_usd": "5.000000",
"reset": "daily",
"spent_usd": "1.200000",
"reserved_usd": "0.020000",
"remaining_usd": "3.780000",
"resets_at": "2026-10-01T00:00:00.000Z"
}
}
```
# Chat Completions (https://docs.binference.io/api/chat-completions)
The OpenAI chat format, for every model.
Send a conversation, get the next message. It's the format most tools speak: the OpenAI SDKs, the Vercel AI SDK, LangChain, OpenClaw and Hermes.
Stream long answers with `stream: true`, and add `stream_options: { "include_usage": true }` to get the cost in the last chunk. See [Streaming](/features/streaming).
`POST https://binference.io/api/v1/chat/completions`
Send your key as `Authorization: Bearer binf_...` or `x-api-key`.
## Request body
- `model` (string, required): A model id from GET /models, such as `anthropic/claude-sonnet-5.5`. Add `:online` for web search, or `:nitro`, `:floor` or `:exacto` to steer which provider runs it.
- `messages` (object[], required): The conversation so far, oldest first.
- `role` ("system" | "user" | "assistant" | "tool", required): Who wrote the message.
- `content` (string | object[], required): Text, or parts: text and image URLs. Send images and files as URLs, not inline data: a request body is at most 4 MB.
- `max_tokens` (integer): The longest answer, in tokens. Set it: a call reserves its longest possible answer while it runs, and without a limit that is the model's whole output. `max_completion_tokens` works too.
- `stream` (boolean): Send the answer as it is written, as server-sent events. Recommended for anything longer than a sentence: streams may run 30 minutes, other calls 13.
- `stream_options` (object): `{ "include_usage": true }` adds a last chunk with `usage`, including what the call cost.
- `temperature` (number): Randomness, 0 to 2. Lower is more predictable.
- `tools` (object[]): Functions the model may call, in the format's own shape. Your code runs them and sends the results back. Web search and web fetch are also served.
- `tool_choice` (string | object): `"auto"`, `"none"`, `"required"` or one function by name.
- `response_format` (object): Structured output: `{ "type": "json_schema", "json_schema": { ... } }` makes the answer match your schema.
- `reasoning` (object): For reasoning models: `{ "effort": "low" | "medium" | "high" }`. Reasoning tokens are billed as output.
- `models` (string[]): Fallback models, tried in order when the first is busy or down. The reserve covers the dearest of them.
## Response
- `id` (string): The answer's id.
- `choices` (object[]): The answer: `message.content`, any `message.tool_calls`, and `finish_reason`.
- `usage` (object): Tokens and what the call cost.
- `prompt_tokens` (integer): Tokens read.
- `completion_tokens` (integer): Tokens written, reasoning included.
- `cost` (number): What this call is charged, in dollars. Exactly what leaves the balance.
## Response headers
- `x-binference-call-id` (string): This call's id on your Calls page. Quote it when you contact us.
- `x-generation-id` (string): The model provider's id for the answer.
- `retry-after` (integer): On a 429 or 503: whole seconds to wait before sending again.
- `retry-after-ms` (integer): The same wait in milliseconds. The OpenAI and Anthropic SDKs read it.
## Example: curl
```bash
curl https://binference.io/api/v1/chat/completions \
-H "Authorization: Bearer $BINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"messages": [
{
"role": "system",
"content": "You are a concise trading assistant."
},
{
"role": "user",
"content": "Summarize the last 24 hours of trades in one line."
}
],
"max_tokens": 256
}'
```
## Example: TypeScript
```ts
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://binference.io/api/v1",
apiKey: process.env.BINF_API_KEY,
});
const reply = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5.5",
messages: [
{
role: "system",
content: "You are a concise trading assistant."
},
{
role: "user",
content: "Summarize the last 24 hours of trades in one line."
}
],
max_tokens: 256
});
console.log(reply.choices[0].message.content);
```
## Example: Python
```python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://binference.io/api/v1",
api_key=os.environ["BINF_API_KEY"],
)
reply = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[
{
"role": "system",
"content": "You are a concise trading assistant.",
},
{
"role": "user",
"content": "Summarize the last 24 hours of trades in one line.",
},
],
max_tokens=256,
)
print(reply.choices[0].message.content)
```
## Example response
```json
{
"id": "gen-1790755652-kQ3hTz",
"object": "chat.completion",
"created": 1790755652,
"model": "anthropic/claude-sonnet-5.5",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "12 trades today: 9 wins, net +3.4 BNB, biggest move on $NOVA."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 31,
"completion_tokens": 22,
"total_tokens": 53,
"cost": 0.000338
}
}
```
## Example stream
```txt
data: {"id":"gen-1790755652-kQ3hTz","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":"12 trades"}}]}
data: {"id":"gen-1790755652-kQ3hTz","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" today: 9 wins"}}]}
data: {"id":"gen-1790755652-kQ3hTz","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":31,"completion_tokens":22,"total_tokens":53,"cost":0.000338}}
data: [DONE]
```
# Errors (https://docs.binference.io/api/errors)
Every error the API returns, why it happens, what to do and whether to retry.
Every error carries a stable `code` in the body, in the format's own shape. Branch on the `code`, not on the message: messages may be reworded.
| Status | Code | Kind | When | What to do | Retry |
| --- | --- | --- | --- | --- | --- |
| 401 | `invalid_api_key` | Key and access | The key is missing, mistyped, unknown or revoked. A key revoked while a call starts is refused at its reserve. | Send the full key as `Authorization: Bearer binf_...` or `x-api-key: binf_...`. If it was revoked, make a new one. | No |
| 403 | `agent_inactive` | Key and access | The agent's launch is not confirmed on chain yet, or its launch fee is unpaid. | Finish the launch on the agent's page. Keys work once it says the agent is live. | No |
| 403 | `agent_suspended` | Key and access | The agent was suspended. None of its keys can start calls. | Contact bInference. | No |
| 400 | `invalid_request` | Request | The body is not JSON, or misses a field the format needs, such as `model` or `messages`. A model's provider can refuse a parameter the same way. | Check the body against the endpoint's reference, and the parameters against the model. | No |
| 400 | `unknown_model` | Request | No listed model has this id. | Copy an id from GET /models or the models page. Ids look like `anthropic/claude-sonnet-5.5`. | No |
| 400 | `model_not_served` | Request | Free models, `:free` and `:batch` versions, retired versions and routers that pick a model after the call. Their price can't be reserved up front. | Use the paid model's own id, or add `:nitro`, `:floor`, `:exacto` or `:online` to it. | No |
| 400 | `tool_not_served` | Request | Tools that run other models, or that keep files, memory or a sandbox on a shared account: image generation, hosted code execution, file search, memory, X search. Also a tool loop with no step limit. | Remove the tool, or run it on your own machine. Web search and web fetch are served. Set `max_tool_calls` on tool loops. | No |
| 413 | `request_too_large` | Request | The body is over 4 MB, or the prompt is longer than the model's context. | Send images and files as URLs, not inline base64. In Claude Code, run /compact or /clear. | No |
| 400 | `request_over_capacity` | Request | One request's reserve is larger than what the gateway can hold for a single call at the moment. | Lower `max_tokens` (`max_output_tokens` on Responses), then send it again. | No |
| 404 | `not_found` | Request | The path doesn't exist, such as embeddings or token counting. | Use one of the five endpoints in this reference. | No |
| 402 | `insufficient_balance` | Balance and limits | What the agent can spend now is less than this call's reserve, its worst-case cost. The message names both amounts. | Lower `max_tokens` so the reserve fits, wait for running calls to finish or add credit. Agents earn credit from their token's trading fees. | No |
| 402 | `key_limit` | Balance and limits | The key's own daily, weekly or monthly limit can't cover this call's reserve. The agent's other keys keep working. | Wait for the reset named in the message, lower `max_tokens` or raise the limit on the keys page. | No |
| 429 | `rate_limited` | Pace | The agent started more than 60 calls at once, or more than 600 a minute. | Wait for Retry-After-Ms and send again. Official OpenAI and Anthropic SDKs do this for you. | After Retry-After-Ms |
| 429 | `too_many_running_calls` | Pace | The agent has 8 calls running. A slot frees the moment one ends. | Queue calls on your side, or wait the 2 to 8 seconds the headers name. | After Retry-After-Ms |
| 429 | `daily_share_used` | Pace | Rare. On a day when 80% of the shared capacity is spent, an agent that already used its fair share waits, so the rest stays for everyone. | Wait until 00:00 UTC. It often lifts sooner, when capacity is added. | After Retry-After-Ms |
| 503/529 | `model_busy` | Busy or down | Every provider of this model is at capacity. On /messages it is 529 `overloaded_error`, so Claude Code moves to its fallback model. | Wait for Retry-After-Ms, or use another model. Listing fallbacks in `models` lets a call run on the next one. | After Retry-After-Ms |
| 503 | `gateway_busy` | Busy or down | The gateway is at its capacity for a moment. | Wait for Retry-After-Ms and send again. | After Retry-After-Ms |
| 503 | `gateway_capacity` | Busy or down | The gateway is topping up its capacity. bInference is alerted at once. | Wait for Retry-After-Ms and send again. | After Retry-After-Ms |
| 503 | `gateway_unavailable` | Busy or down | The gateway can't reach its models for a moment. bInference is alerted at once. | Wait for Retry-After-Ms and send again. | After Retry-After-Ms |
| 503 | `catalog_unavailable` | Busy or down | The gateway couldn't read model prices, so it can't size the reserve. | Send again after 5 seconds. | After Retry-After-Ms |
| 503 | `service_unavailable` | Busy or down | The gateway couldn't reach its database before the reserve. Nothing was reserved. | Send again after about 2 seconds. | After Retry-After-Ms |
| 500 | `gateway_error` | Busy or down | A fault in the gateway after the reserve. The reserve is released and the call is free. | Send it again. If it keeps failing, contact bInference with the `x-binference-call-id` header. | Yes |
| 504 | `upstream_timeout` | Model provider | The model sent nothing before the time limit: 30 minutes streamed, 13 minutes not streamed. | Stream long calls, lower `max_tokens` or use a faster model. | Yes |
| 502 | `upstream_unreachable` | Model provider | The connection to the model's provider failed. | Send it again. | Yes |
| 502 | `provider_error` | Model provider | The model's provider failed, or refused this one call. | Send it again, or use another model. | Yes |
| 503 | `provider_unavailable` | Model provider | The model is down at its providers. | Send it again shortly, or use another model. | Yes |
| 404 | `no_provider` | Model provider | No provider of the model supports every option in the request, such as tools, images or a long context. | Drop an option, or pick a model that lists it. | No |
| 403 | `provider_refused` | Model provider | The provider's own policy refused the request. | Change the request, or use another model. | No |
| 422 | `unprocessable` | Model provider | The provider accepted the request but couldn't run it. | Check the parameters against the model, or use another model. | No |
| 408 | `timeout` | Model provider | The provider gave up on the call (408 or 504). | Send it again, stream it or lower `max_tokens`. | Yes |
| 502 | `upstream_error` | Model provider | Any other error from the provider. | Send it again shortly. | Yes |
## The short version [#the-short-version]
| Status | Means | Retry? |
| ----------- | ------------------------------------------------------ | ------------------------------------ |
| `400` | The request can't be served as sent | No. Fix it first |
| `401` | The key is missing, wrong or revoked | No |
| `402` | The balance or the key's limit can't cover the reserve | No. Lower `max_tokens` or add credit |
| `403` | The agent isn't active | No |
| `404` | Unknown path, or no provider for this request | No |
| `413` | The body is over 4 MB | No. Send files as URLs |
| `429` | Too many running calls, or too fast | Yes, after `Retry-After-Ms` |
| `500` | The gateway failed. The call is free | Yes |
| `502` `504` | The model's provider failed | Yes |
| `503` `529` | Busy for a moment | Yes, after `Retry-After-Ms` |
## Errors inside a stream [#errors-inside-a-stream]
Once an answer has started streaming, an error can't change the status any more. It arrives as the format's own error event, so SDKs raise it instead of hanging:
Chat Completions
Responses
Messages
```txt
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"error"}],"error":{"message":"The model took too long to answer. Try again.","type":"server_error","code":"timeout"}}
```
```txt
event: error
data: {"type":"error","code":"timeout","message":"The model took too long to answer. Try again.","param":null}
```
```txt
event: error
data: {"type":"error","error":{"type":"api_error","message":"The model took too long to answer. Try again."}}
```
The call is charged for what was written before the error.
# API overview (https://docs.binference.io/api)
One base URL, one key, three formats. Everything you need before the first call.
## Base URL [#base-url]
```txt
https://binference.io/api/v1
```
For the Anthropic SDKs and Claude Code, use `https://binference.io/api`: they add `/v1` themselves.
## Endpoints [#endpoints]
| Method | Path | Format | Key |
| ------ | -------------------------------------------- | ------------------------- | --- |
| `POST` | [`/chat/completions`](/api/chat-completions) | OpenAI Chat Completions | Yes |
| `POST` | [`/responses`](/api/responses) | OpenAI Responses | Yes |
| `POST` | [`/messages`](/api/messages) | Anthropic Messages | Yes |
| `GET` | [`/models`](/api/models) | Every model and its price | No |
| `GET` | [`/balance`](/api/balance) | What the key can spend | Yes |
Any other path under `/api/v1`, like embeddings, answers `404 not_found` in JSON.
## Authentication [#authentication]
Send your `binf_` key as a bearer token or as `x-api-key`:
```http
Authorization: Bearer binf_...
```
A missing, wrong or revoked key answers `401 invalid_api_key`. See [API keys](/keys).
## Requests [#requests]
* **JSON bodies,** up to 4 MB. Send images and files as URLs, not inline data.
* **Bodies pass through as sent.** Every field of each format works, beyond the ones these pages list.
* **Stateless.** Nothing is stored between calls. Send the conversation every time.
* **From anywhere.** The API accepts calls from any origin, so browser tools work. Keys in web pages can be read by visitors: keep them on a server.
## Answers [#answers]
Answers follow each format's own shape, with three additions:
| Field or header | What it is |
| ---------------------- | ------------------------------------------------------------------------- |
| `usage.cost` | What the call was charged, in dollars. On Chat Completions and Responses. |
| `x-binference-call-id` | The call's id on your Calls page. Quote it when you contact us. |
| `x-generation-id` | The provider's id for the answer. |
## Errors [#errors]
Errors use each format's own shape, with a stable `code` to branch on:
Chat Completions and Responses
Messages
```json
{
"error": {
"message": "This agent is starting calls faster than 600 a minute. Retry in 1 second.",
"type": "rate_limit_error",
"code": "rate_limited",
"param": null
}
}
```
```json
{
"type": "error",
"error": {
"type": "rate_limit_error",
"message": "This agent is starting calls faster than 600 a minute. Retry in 1 second.",
"code": "rate_limited"
}
}
```
See every code in [Errors](/api/errors), and the limits in [Rate limits](/api/rate-limits).
# Messages (https://docs.binference.io/api/messages)
The Anthropic Messages format, for Claude Code and every model.
Anthropic's format, used by Claude Code and the Anthropic SDKs. It works with every listed model, not only Claude.
`max_tokens` is required, as in Anthropic's own API. A busy model answers `529 overloaded_error`, so Claude Code moves to its fallback model. Messages answers carry no cost: find each call's charge on your [Calls page](https://binference.io/account/calls).
`POST https://binference.io/api/v1/messages`
Send your key as `Authorization: Bearer binf_...` or `x-api-key`.
## Request body
- `model` (string, required): A model id from GET /models. Claude clients' own names work too: `claude-sonnet-5-5`, dated names and the `[1m]` suffix map to the listed model.
- `messages` (object[], required): The conversation so far, alternating `user` and `assistant`.
- `max_tokens` (integer, required): The longest answer, in tokens. The call reserves this much output while it runs.
- `system` (string | object[]): The system prompt.
- `stream` (boolean): Send the answer as it is written, as server-sent events. Recommended for anything longer than a sentence: streams may run 30 minutes, other calls 13.
- `temperature` (number): Randomness, 0 to 2. Lower is more predictable.
- `tools` (object[]): Functions the model may call, in the format's own shape. Your code runs them and sends the results back. Web search and web fetch are also served.
- `thinking` (object): `{ "type": "enabled", "budget_tokens": 2048 }` for extended thinking. Thinking is billed as output.
## Headers
- `anthropic-version` (string): Passed on as sent. The Anthropic SDKs set it.
- `anthropic-beta` (string): Passed on as sent, for beta features.
## Response
- `id` (string): The message's id.
- `content` (object[]): Blocks: `text`, `thinking` and `tool_use`.
- `stop_reason` (string): Why the answer ended.
- `usage` (object): `input_tokens` and `output_tokens`. The Messages format carries no cost: find each call's charge on your Calls page within a minute.
## Response headers
- `x-binference-call-id` (string): This call's id on your Calls page. Quote it when you contact us.
- `x-generation-id` (string): The model provider's id for the answer.
- `retry-after` (integer): On a 429 or 503: whole seconds to wait before sending again.
- `retry-after-ms` (integer): The same wait in milliseconds. The OpenAI and Anthropic SDKs read it.
## Example: curl
```bash
curl https://binference.io/api/v1/messages \
-H "Authorization: Bearer $BINF_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5.5",
"max_tokens": 256,
"system": "You are a concise trading assistant.",
"messages": [
{
"role": "user",
"content": "Summarize the last 24 hours of trades in one line."
}
]
}'
```
## Example: TypeScript
```ts
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: "https://binference.io/api",
apiKey: process.env.BINF_API_KEY,
});
const message = await client.messages.create({
model: "anthropic/claude-sonnet-5.5",
max_tokens: 256,
system: "You are a concise trading assistant.",
messages: [
{
role: "user",
content: "Summarize the last 24 hours of trades in one line."
}
]
});
console.log(message.content);
```
## Example: Python
```python
import os
from anthropic import Anthropic
client = Anthropic(
base_url="https://binference.io/api",
api_key=os.environ["BINF_API_KEY"],
)
message = client.messages.create(
model="anthropic/claude-sonnet-5.5",
max_tokens=256,
system="You are a concise trading assistant.",
messages=[
{
"role": "user",
"content": "Summarize the last 24 hours of trades in one line.",
},
],
)
print(message.content)
```
## Example response
```json
{
"id": "msg_01XFDUDYJgAACzvnptvVoYEL",
"type": "message",
"role": "assistant",
"model": "anthropic/claude-sonnet-5.5",
"content": [
{
"type": "text",
"text": "12 trades today: 9 wins, net +3.4 BNB, biggest move on $NOVA."
}
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 27,
"output_tokens": 22
}
}
```
## Example stream
```txt
event: message_start
data: {"type":"message_start","message":{"id":"msg_01XFDUDYJgAACzvnptvVoYEL","role":"assistant","usage":{"input_tokens":27,"output_tokens":1}}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"12 trades today"}}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":22}}
event: message_stop
data: {"type":"message_stop"}
```
# List models (https://docs.binference.io/api/models)
Every model you can call and its price. No key needed.
The same list the gateway serves: every model with a fixed price, and the prices you pay. Prices are in dollars as decimal strings, per token unless named otherwise.
It changes when models arrive or leave, at most every hour, so cache it for a few minutes.
`GET https://binference.io/api/v1/models`
No key needed.
## Response
- `data` (object[]): One entry per model.
- `id` (string): The id to send as `model`.
- `name` (string): Its display name.
- `context_length` (integer): The longest prompt plus answer, in tokens.
- `max_output_tokens` (integer | null): The longest answer it can write.
- `pricing` (object): What you pay, in dollars as decimal strings: `prompt`, `completion` and `reasoning` per token, `request` per call, `image` per image, `web_search` per search.
## Example: curl
```bash
curl https://binference.io/api/v1/models
```
## Example: TypeScript
```ts
const response = await fetch("https://binference.io/api/v1/models");
const models = await response.json();
```
## Example: Python
```python
import os
import requests
response = requests.get("https://binference.io/api/v1/models")
print(response.json())
```
## Example response
```json
{
"object": "list",
"data": [
{
"id": "anthropic/claude-sonnet-5.5",
"object": "model",
"owned_by": "anthropic",
"name": "Anthropic: Claude Sonnet 5.5",
"context_length": 1000000,
"max_output_tokens": 128000,
"pricing": {
"prompt": "0.0000024",
"completion": "0.000012",
"reasoning": "0.000012",
"request": "0",
"image": "0",
"web_search": "0.012"
}
}
]
}
```
# Playground (https://docs.binference.io/api/playground)
Try any model with your key. See the reserve before you send, then watch the answer stream in.
Try the API in the browser at https://docs.binference.io/api/playground.
# Rate limits (https://docs.binference.io/api/rate-limits)
How many calls an agent may run and start, how to read the waits and how to stay under.
Limits are per agent (or per account), not per key: more keys don't add room. A refused call is always free and never counts against the limits.
## Try it [#try-it]
## The two limits you can meet [#the-two-limits-you-can-meet]
**8 calls running at once.** The 9th is refused with `429 too_many_running_calls` and a wait of 2 to 8 seconds, spread so a refused batch doesn't return all at once. A slot frees the moment a call ends.
**600 calls started a minute.** Starts are spaced on a schedule: one every 100 ms, with up to 60 at once after a quiet spell. Faster than that is refused with `429 rate_limited` and the exact wait.
With calls that take a second or more, the 8-call limit is the one you meet first. The rate matters for many very short calls.
## Reading the headers [#reading-the-headers]
Every `429` and `503` names its wait:
```http
HTTP/1.1 429 Too Many Requests
retry-after: 2
retry-after-ms: 1180
```
* `retry-after-ms` is exact; `retry-after` is the same in whole seconds.
* Waits run from 1 second to 60. We add up to a fifth at random so clients told the same wait don't all come back at once.
* The OpenAI and Anthropic SDKs read these headers and wait on their own.
## Busy days [#busy-days]
Every agent shares the day's AI capacity. When 80% of it is used, an agent that already had its fair share today waits until 00:00 UTC with `429 daily_share_used`, so the rest stays for everyone. It's rare, and it lifts sooner when capacity is added.
## Stay under [#stay-under]
* **Queue on your side.** Run at most 8 calls per agent at once. A simple semaphore does it.
* **Stream.** Streamed calls hold a slot as long as other calls, but fail less on slow networks.
* **Let the SDK retry.** It already honors `Retry-After-Ms`. See [Retries](/api/retries).
# Responses (https://docs.binference.io/api/responses)
The OpenAI Responses format, which Codex speaks.
The newer OpenAI format, with `input` and `instructions` in place of messages. Codex uses it.
It is stateless here: nothing is stored, so `previous_response_id` and saved prompts aren't served. Send the conversation as `input` items each time.
`POST https://binference.io/api/v1/responses`
Send your key as `Authorization: Bearer binf_...` or `x-api-key`.
## Request body
- `model` (string, required): A model id from GET /models, such as `anthropic/claude-sonnet-5.5`. Add `:online` for web search, or `:nitro`, `:floor` or `:exacto` to steer which provider runs it.
- `input` (string | object[], required): A prompt, or the conversation as input items.
- `instructions` (string): The system prompt.
- `max_output_tokens` (integer): The longest answer, in tokens. Set it: the call reserves its longest possible answer while it runs.
- `stream` (boolean): Send the answer as it is written, as server-sent events. Recommended for anything longer than a sentence: streams may run 30 minutes, other calls 13.
- `temperature` (number): Randomness, 0 to 2. Lower is more predictable.
- `tools` (object[]): Functions the model may call, in the format's own shape. Your code runs them and sends the results back. Web search and web fetch are also served.
- `max_tool_calls` (integer): The most tool steps in one call. Web search loops reserve for every step, so a lower number reserves less.
- `reasoning` (object): For reasoning models: `{ "effort": "low" | "medium" | "high" }`.
- `models` (string[]): Fallback models, tried in order when the first is busy or down. The reserve covers the dearest of them.
## Response
- `id` (string): The response's id.
- `output` (object[]): Output items: messages, reasoning and tool calls. `output_text` joins the text.
- `usage` (object): `input_tokens`, `output_tokens` and `cost`: what this call is charged, in dollars.
## Response headers
- `x-binference-call-id` (string): This call's id on your Calls page. Quote it when you contact us.
- `x-generation-id` (string): The model provider's id for the answer.
- `retry-after` (integer): On a 429 or 503: whole seconds to wait before sending again.
- `retry-after-ms` (integer): The same wait in milliseconds. The OpenAI and Anthropic SDKs read it.
## Example: curl
```bash
curl https://binference.io/api/v1/responses \
-H "Authorization: Bearer $BINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-6.1-sol",
"instructions": "You are a concise trading assistant.",
"input": "Summarize the last 24 hours of trades in one line.",
"max_output_tokens": 256
}'
```
## Example: TypeScript
```ts
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://binference.io/api/v1",
apiKey: process.env.BINF_API_KEY,
});
const reply = await client.responses.create({
model: "openai/gpt-6.1-sol",
instructions: "You are a concise trading assistant.",
input: "Summarize the last 24 hours of trades in one line.",
max_output_tokens: 256
});
console.log(reply.output_text);
```
## Example: Python
```python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://binference.io/api/v1",
api_key=os.environ["BINF_API_KEY"],
)
reply = client.responses.create(
model="openai/gpt-6.1-sol",
instructions="You are a concise trading assistant.",
input="Summarize the last 24 hours of trades in one line.",
max_output_tokens=256,
)
print(reply.output_text)
```
## Example response
```json
{
"id": "resp_7Yq2c1",
"object": "response",
"status": "completed",
"model": "openai/gpt-6.1-sol",
"output": [
{
"type": "message",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "12 trades today: 9 wins, net +3.4 BNB, biggest move on $NOVA."
}
]
}
],
"usage": {
"input_tokens": 29,
"output_tokens": 21,
"total_tokens": 50,
"cost": 0.000322
}
}
```
## Example stream
```txt
event: response.created
data: {"type":"response.created","response":{"id":"resp_7Yq2c1","status":"in_progress"}}
event: response.output_text.delta
data: {"type":"response.output_text.delta","delta":"12 trades today"}
event: response.completed
data: {"type":"response.completed","response":{"id":"resp_7Yq2c1","status":"completed","usage":{"input_tokens":29,"output_tokens":21,"total_tokens":50,"cost":0.000322}}}
```
# Retries (https://docs.binference.io/api/retries)
Which errors to retry, how long to wait and code that does it right.
## What to retry [#what-to-retry]
| Answer | Retry? | How |
| ---------------------------------------- | ------- | --------------------------------------------------------- |
| `429`, `503`, `529` | Yes | After `Retry-After-Ms` |
| `500`, `502`, `504` | Yes | After a short backoff, 1 to 2 seconds |
| `400`, `401`, `402`, `403`, `404`, `413` | No | The same call fails the same way. Fix it first |
| Network error before any answer | Yes | After a short backoff |
| Error inside a stream | Usually | Resend the whole call. You paid only for what was written |
The gateway never retries a call itself: your SDK does, so tries don't multiply.
## The official SDKs already do it [#the-official-sdks-already-do-it]
The OpenAI and Anthropic SDKs retry `429` and `5xx` answers twice by default and wait for `Retry-After-Ms`. For long unattended jobs, raise it:
TypeScript
Python
```ts
const client = new OpenAI({
baseURL: "https://binference.io/api/v1",
apiKey: process.env.BINF_API_KEY,
maxRetries: 5,
});
```
```python
client = OpenAI(
base_url="https://binference.io/api/v1",
api_key=os.environ["BINF_API_KEY"],
max_retries=5,
)
```
## Doing it yourself [#doing-it-yourself]
With plain `fetch`, read the wait from the headers and add a little randomness:
```ts
async function call(body: unknown, tries = 5): Promise {
for (let attempt = 1; ; attempt++) {
const response = await fetch("https://binference.io/api/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.BINF_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
});
const retryable = response.status === 429 || response.status >= 500;
if (!retryable || attempt >= tries) return response;
const named = Number(response.headers.get("retry-after-ms"));
const wait = Number.isFinite(named) && named > 0 ? named : 1_000 * 2 ** (attempt - 1);
await new Promise((resolve) => setTimeout(resolve, wait + Math.random() * 250));
}
}
```
## Time limits [#time-limits]
* **Streamed calls:** up to 30 minutes.
* **Calls that aren't streamed:** up to 13 minutes.
* At the limit, the call ends with a clean error and is charged for what was written.
Set your client's timeout above these, or stream: some SDKs give up on a silent connection after 10 minutes.