MCP server
Let your agent check its own credit, usage and limits, manage its own keys and upload images.
Connect the bInference MCP server and your agent can answer "how much AI budget do I have left?" on its own, make a key for a new worker, or revoke one it no longer needs. It signs in with the binf_ key it already uses for model calls.
https://binference.io/api/mcpWhat a key can do
Every tool acts on the agent, or the account, that owns the key. A key never sees another agent.
| Key with a spending limit | Key without a limit | |
|---|---|---|
| Read credit, usage, calls, limits and keys | Yes | Yes |
| Upload images for models to read | Yes | Yes |
| Create, change and revoke keys | No | Yes |
A key with a limit only reads, so it can never make itself a key without one. A key without a limit can already spend the whole balance, so managing keys gives it nothing more to spend.
Which key to connect
If your agent only needs to watch its budget, connect a key with a daily limit. Connect a key without a limit only to an agent you trust to manage keys.
Connect it
claude mcp add --transport http binference https://binference.io/api/mcp --header "Authorization: Bearer binf_..."Add --scope user to use it in every project. Then run /mcp in Claude Code: binference shows as connected.
Any other client works the same way if it can connect to a Streamable HTTP server and send a header. Apps that can only connect through a browser sign-in (OAuth), like Claude Desktop, ChatGPT, Grok and Muse, can't connect yet.
From your own code
Use the official MCP SDK: npm install @modelcontextprotocol/client, or pip install "mcp>=2.2" for Python.
import { Client, StreamableHTTPClientTransport } from "@modelcontextprotocol/client";
const client = new Client({ name: "my-agent", version: "1.0.0" });
await client.connect(
new StreamableHTTPClientTransport(new URL("https://binference.io/api/mcp"), {
requestInit: { headers: { Authorization: `Bearer ${process.env.BINF_API_KEY}` } },
}),
);
const balance = await client.callTool({ name: "get_balance", arguments: {} });
console.log(balance.structuredContent);Try it
Ask your agent:
How much bInference credit do I have, and how much of it expires this week?It calls get_balance and answers from what it can spend now and what expires on each day.
| You ask | The agent uses |
|---|---|
| "What did I spend on AI this week, and on which models?" | get_usage |
| "Why did my last few calls fail?" | list_calls |
| "How close is this key to its daily limit?" | get_limits |
| "Make a key called worker-2 with a $5 daily limit." | create_key |
| "Revoke the key called laptop." | list_keys, then revoke_key |
| "Describe the screenshot at ./bug.png." | create_upload, then a model |
The tools
| Tool | What it does | Needs a key without a limit |
|---|---|---|
get_balance | What new calls can spend now, what running calls hold, what expires when | No |
get_limits | The key's spending limit and the agent's call limits | No |
get_usage | Spending by day and by model, up to 90 days back | No |
list_calls | Recent calls with model, outcome, error code, tokens and charge | No |
list_keys | Active keys with their limits and this month's spending | No |
create_key | A new key, with or without a spending limit | Yes |
update_key | Rename a key, or set, change or remove its limit | Yes |
revoke_key | Revoke a key | Yes |
create_upload | A link to upload a local image to, and the link to send models | No |
Every input, output and refusal is in MCP tools.
Limits
| Limit | How much |
|---|---|
| MCP requests | 60 a minute per key, up to 20 at once |
| Keys an agent makes over MCP | 20 in any 24 hours |
| Active keys | 10 per agent, and 10 for your account |
| Image uploads | 200 per agent in any 24 hours |
A key over its request rate gets 429 with the wait in retry-after. The limits are per key, so one busy key never slows the agent's others. Model calls have their own rate limits. get_limits returns all of these.
Keep it safe
- A new key is shown in the conversation.
create_keyreturns the key once, and the agent's model sees it. Ask the agent to put it straight into an environment variable or a secret store, and don't create keys in a conversation you share. - The key is checked on every request. A key you revoke on the API keys page stops at once, here too.
- You see everything the agent does. Keys it makes show up on your API keys page marked "over MCP", with the key that made them, and you can limit or revoke them there.
- The tools are free. They spend no credit. Only model calls do. Uploads need an agent with credit, since they're for sending images to models.