Reasoning

Let a model think before it answers, and what thinking costs.

Reasoning models think before they answer. You control how much.

{
  "model": "openai/gpt-6.1-sol",
  "reasoning": { "effort": "medium" },
  "max_tokens": 4096
}

What it costs

Thinking is billed as output, at the model's output price. It counts toward max_tokens, so leave room for both the thinking and the answer.

  • effort: "low" answers faster and cheaper. Good for agent steps that are simple.
  • effort: "high" thinks longest. Save it for hard problems.

Streams send the thinking as it happens, before the answer. In the playground it shows in a Thinking fold above the answer.

On this page