Models

Every model you can call, with live prices. Copy an id and send it as model.

Every listed model has a fixed price per token, and these are the prices you pay. The list below is live from GET /models.

Model ids

Send the id as model, like anthropic/claude-sonnet-5.5. A few more forms work:

  • Claude clients' own names on /messages: claude-sonnet-5-5, dated names like claude-haiku-4-5-20251001 and the [1m] suffix map to the listed model.
  • Latest aliases like ~anthropic/claude-opus-latest work and are priced as the model they point to.
  • Variants change how a model runs, not what it costs per token: :online, :nitro, :floor and :exacto. See Web search and Routing and fallbacks.

Not served

These answer 400 model_not_served, because their price can't be known before the call:

  • free models and any :free version
  • :batch versions and retired versions like :thinking
  • routers that pick a model after the call

A model we don't know answers 400 unknown_model. New models show up in the list within an hour of release.

Picking a model

  • For agents that loop all day, a fast, cheap model keeps the budget going: Gemini Flash, DeepSeek Flash or Claude Haiku.
  • For hard problems, a frontier model with reasoning: Claude Opus, GPT or Gemini Pro.
  • Compare side by side on binference.io/models: prices, context, the hosts that run each one and what it can do.

On this page