# Quickstart (https://docs.binference.io/quickstart)

Your first call in two minutes. If you've used the OpenAI or Anthropic API, you already know how.



<Callout title="Using Binance Agent OS?">
  Follow the [Binance Agent OS guide](/agent-os): it sets up the model and the Binance
  connection together.
</Callout>

<Steps>
  <Step>
    ### Get a key [#get-a-key]

    Sign in on [binference.io](https://binference.io/account/keys) with your wallet and open **API keys**.

    * **Agent key:** spends an agent's own budget, filled by its token's trading fees. You need to have launched the agent.
    * **Account key:** spends credit you buy with BNB, for any use.

    The key is shown once. It starts with `binf_`. Keep it in an environment variable:

    ```bash
    export BINF_API_KEY="binf_..."
    ```
  </Step>

  <Step>
    ### Make a call [#make-a-call]

    Change two settings in the client you already use: the address and the key. That's the only change.

    <CodeBlockTabs defaultValue="curl">
      <CodeBlockTabsList>
        <CodeBlockTabsTrigger value="curl">
          curl
        </CodeBlockTabsTrigger>

        <CodeBlockTabsTrigger value="TypeScript">
          TypeScript
        </CodeBlockTabsTrigger>

        <CodeBlockTabsTrigger value="Python">
          Python
        </CodeBlockTabsTrigger>
      </CodeBlockTabsList>

      <CodeBlockTab value="curl">
        ```bash
        curl https://binference.io/api/v1/chat/completions \
          -H "Authorization: Bearer $BINF_API_KEY" \
          -H "Content-Type: application/json" \
          -d '{
            "model": "anthropic/claude-sonnet-5.5",
            "messages": [{ "role": "user", "content": "Say hello in five words." }],
            "max_tokens": 256
          }'
        ```
      </CodeBlockTab>

      <CodeBlockTab value="TypeScript">
        ```ts
        import OpenAI from "openai";

        const client = new OpenAI({
          baseURL: "https://binference.io/api/v1",
          apiKey: process.env.BINF_API_KEY,
        });

        const reply = await client.chat.completions.create({
          model: "anthropic/claude-sonnet-5.5",
          messages: [{ role: "user", content: "Say hello in five words." }],
          max_tokens: 256,
        });

        console.log(reply.choices[0].message.content);
        ```
      </CodeBlockTab>

      <CodeBlockTab value="Python">
        ```python
        import os
        from openai import OpenAI

        client = OpenAI(
            base_url="https://binference.io/api/v1",
            api_key=os.environ["BINF_API_KEY"],
        )

        reply = client.chat.completions.create(
            model="anthropic/claude-sonnet-5.5",
            messages=[{"role": "user", "content": "Say hello in five words."}],
            max_tokens=256,
        )

        print(reply.choices[0].message.content)
        ```
      </CodeBlockTab>
    </CodeBlockTabs>
  </Step>

  <Step>
    ### Check what it cost [#check-what-it-cost]

    Every answer's `usage.cost` is what the call was charged, in dollars. Your balance is one call away:

    ```bash
    curl https://binference.io/api/v1/balance \
      -H "Authorization: Bearer $BINF_API_KEY"
    ```
  </Step>
</Steps>

<Callout title="Always set max_tokens">
  While a call runs it reserves its longest possible answer from your balance. Without a
  limit that is the model's whole output, which can be dollars. You pay only what the
  model writes, and the rest comes back when it ends. [How a call is
  paid](/how-it-works/calls)
</Callout>

## Next [#next]

<Cards>
  <Card title="Try it in the playground" href="/api/playground" description="Paste your key and watch answers stream in." />

  <Card title="Connect your tools" href="/clients/claude-code" description="Claude Code, Codex, OpenClaw, Hermes and the SDKs." />

  <Card title="Pick a model" href="/models" description="Every model and its price per million tokens." />

  <Card title="Handle errors" href="/api/errors" description="What each error means and whether to retry." />
</Cards>
