OpenClaw

Add bInference as a custom provider in OpenClaw.

Add a provider to ~/.openclaw/openclaw.json, then allow its model for your agents:

~/.openclaw/openclaw.json
{
  models: {
    providers: {
      binference: {
        baseUrl: "https://binference.io/api/v1",
        apiKey: "${BINF_API_KEY}",
        api: "openai-completions",
        models: [
          {
            id: "anthropic/claude-sonnet-5.5",
            name: "Claude Sonnet 5.5",
            contextWindow: 1000000,
            // The longest answer per call. It is also what each call reserves.
            maxTokens: 8192,
          },
        ],
      },
    },
  },
  agents: {
    defaults: {
      model: { primary: "binference/anthropic/claude-sonnet-5.5" },
      models: { "binference/anthropic/claude-sonnet-5.5": { alias: "Sonnet" } },
    },
  },
}

Model ids keep their maker prefix, so the full name is binference/ plus the id from Models.

Keep maxTokens modest

OpenClaw sends maxTokens as each call's output limit, and every call reserves that much from the budget while it runs. 8,192 suits most agent steps.

To use the Messages format instead, set api: "anthropic-messages" and baseUrl: "https://binference.io/api".