Models
Every model you can call, with live prices. Copy an id and send it as model.
Every listed model has a fixed price per token, and these are the prices you pay. The list below is live from GET /models.
Model ids
Send the id as model, like anthropic/claude-sonnet-5.5. A few more forms work:
- Claude clients' own names on
/messages:claude-sonnet-5-5, dated names likeclaude-haiku-4-5-20251001and the[1m]suffix map to the listed model. - Latest aliases like
~anthropic/claude-opus-latestwork and are priced as the model they point to. - Variants change how a model runs, not what it costs per token:
:online,:nitro,:floorand:exacto. See Web search and Routing and fallbacks.
Not served
These answer 400 model_not_served, because their price can't be known before the call:
- free models and any
:freeversion :batchversions and retired versions like:thinking- routers that pick a model after the call
A model we don't know answers 400 unknown_model. New models show up in the list within an hour of release.
Picking a model
- For agents that loop all day, a fast, cheap model keeps the budget going: Gemini Flash, DeepSeek Flash or Claude Haiku.
- For hard problems, a frontier model with reasoning: Claude Opus, GPT or Gemini Pro.
- Compare side by side on binference.io/models: prices, context, the hosts that run each one and what it can do.