Streaming

Get the answer as it's written, and why long calls should always stream.

Add "stream": true and the answer arrives as server-sent events, a few words at a time. Each format streams its own way, exactly as its official API does, so the SDKs read it as usual.

Why stream

  • Longer limit. A streamed call may run 30 minutes. One that isn't streamed stops at 13.
  • Steady connection. Some networks drop a connection that stays silent for minutes. A stream keeps it alive.
  • Faster feel. People see the first words in well under a second on fast models.

Claude Code, Codex, OpenClaw and Hermes stream by default.

The cost of a streamed call

On Chat Completions, add stream_options to get a last chunk with usage, including what the call cost:

{
  "model": "google/gemini-3.8-flash",
  "messages": [{ "role": "user", "content": "Hello" }],
  "max_tokens": 256,
  "stream": true,
  "stream_options": { "include_usage": true }
}

The final event before data: [DONE]:

{
  "choices": [],
  "usage": { "prompt_tokens": 9, "completion_tokens": 12, "cost": 0.000062 }
}

On Responses, the response.completed event carries usage.cost. Messages streams carry token counts but no cost; the charge is on your Calls page within a minute.

When a stream breaks

  • An error after words have arrived comes as the format's own error event, so the SDK raises it instead of hanging. The call is charged for what was written.
  • If you disconnect, the model stops and you're charged for what it wrote before. We never charge a disconnect as zero, because some providers bill the work already done.
  • At the time limit, the stream ends with a clean error event 5 seconds before the cut.

Try it in the playground: the Raw tab shows each event with its time.

On this page