What it costs
What an Agent OS session spends on AI, and five ways to keep it low.
Trading through Agent OS is between you and Binance. What your agent spends on bInference is its model calls: tokens in and out, at the prices on Models.
Where the tokens go
Every turn of a session sends the model:
- the conversation so far, which grows each turn
- Binance's tool list, so the model knows what it can do
- tool results, like an order book or a balance, read as input
- any skills the question calls for
The model's answer, and any reasoning, is output.
An example session
Twenty turns with Claude Sonnet 5.5 at $2.40 per million tokens in and $12 out, averaging 25,000 tokens in and 800 out per turn:
| Tokens | Cost | |
|---|---|---|
| In | 500,000 | $1.20 |
| Out | 16,000 | $0.19 |
| Session | $1.39 |
The same session on Gemini 3.8 Flash ($0.90 in, $4.50 out) costs about $0.52. Models that cache the conversation cost less on long sessions, and Claude Code asks for caching on every call.
Keep it low
- Pick the model for the job. A fast model for market checks, a frontier model for decisions.
- Compact long sessions.
/compactin Claude Code shrinks what every later turn resends. - Grant fewer scopes. Fewer tools means a shorter tool list on every turn.
- Install the skills you use, not the whole hub.
- Set a daily limit on the key. It can't overspend, whatever the session does.
Paid by trading
An agent key spends the agent's own budget: 63¢ of every tax dollar its token's trades pay. An agent whose token trades $10,000 a day at 3% tax earns about $189 of AI a day, well over a hundred sessions like the one above. See Where the money comes from.