Routing and fallbacks
Steer which provider runs a model, and keep going when one is busy.
Most models run at several providers. By default each call goes to a good one, and moves to another when the first is down.
Variants
Add a variant to a model id to steer that choice:
| Variant | What it does |
|---|---|
:nitro | Picks the fastest providers first. Some charge more, so the reserve covers them. |
:floor | Picks the cheapest providers first. |
:exacto | Picks providers known to handle tool calls most accurately. |
:online | Searches the web first. See Web search. |
{ "model": "deepseek/deepseek-v4.1-flash:nitro" }:free, :batch and retired variants like :thinking are not served.
Fallback models
List other models in models, and a call moves down the list when the first is busy or down:
{
"model": "anthropic/claude-sonnet-5.5",
"models": [
"anthropic/claude-sonnet-5.5",
"openai/gpt-6.1-sol",
"google/gemini-3.8-flash"
],
"max_tokens": 1024
}The reserve covers the priciest model on the list. The answer's model field says which one ran.
When every provider is busy
A model busy at all its providers answers 503 model_busy (529 on Messages) with the wait in Retry-After-Ms. That model pauses for the wait, so calls to it get the same answer at once, with nothing reserved. Models on your fallback list keep working.