Images

Send images to any model that reads them, and make images with every image model.

Images work both ways on the same key and balance: models read the images you send, and image models make new ones.

Send images to a model

More than half the models read images. Send one as part of a message, in each format's own shape:

{ "type": "image_url", "image_url": { "url": "https://example.com/chart.png" } }
  • Links or inline. An https link, or the image itself as a data:image/png;base64,... URL (a base64 source on Messages). A request body is at most 4 MB and base64 makes a file a third bigger, so send anything large by link: upload it if it isn't online.
  • Types. PNG, JPEG, WebP and GIF. Any other inline type, such as HEIC, SVG or TIFF, answers 400 image_type_not_served.
  • Which models. Those with image in architecture.input_modalities in GET /models. Images sent to a model that can't read them, with no fallback that can, answer 400 model_no_image_input before anything is reserved.
  • Cost. Images are read as input tokens at the model's price, plus its per-image price (pricing.image) where it has one. The reserve counts 8,000 tokens for each image.

For an image that's too big to send inline, or that only exists on your machine, ask for an upload link, send the image straight to storage, then give the model its link:

# 1. A link to upload to, for this type and this exact size.
curl https://binference.io/api/v1/uploads \
  -H "Authorization: Bearer $BINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{\"content_type\": \"image/png\", \"size\": $(wc -c < photo.png)}"

# 2. The image's bytes, with the Content-Type you named.
curl -X PUT "$UPLOAD_URL" -H "Content-Type: image/png" --data-binary @photo.png

Then send url from the first answer as the image, in any format, to any model that reads images. The image goes straight to storage, never through the API, so the 4 MB body limit doesn't apply.

  • Up to 20 MB an image: PNG, JPEG, WebP or GIF. The upload must be the type and the exact size you asked for, or storage refuses it.
  • Timing. The upload link works for 10 minutes. The image's link works for 24 hours, then the image is deleted.
  • Private. Each image sits under your agent with a random name, and only its own signed link reads it.
  • Free, up to 200 uploads per agent a day, for an agent with credit to spend.
  • Over MCP, agents can do it themselves with create_upload.

Make images

POST /images makes images from a prompt with any image model: GPT Image, Gemini, FLUX, Recraft, Seedream, Qwen, Grok and more. GET /images/models lists them, the fields each one takes and its price.

curl https://binference.io/api/v1/images \
  -H "Authorization: Bearer $BINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "recraft/recraft-v4.1-flash",
    "prompt": "A small red lighthouse on a rock at dusk, flat illustration"
  }'

Each image comes back as base64 in data[].b64_json, with its media_type, and usage.cost says what the call was charged.

  • More than one. n makes several in one call, as many as the model allows (parameters.n).
  • Edit or follow an image. Send it in input_references, as a link or inline, on models with reads_images: true.
  • Stream. stream: true sends partial images as they form, then the finished one, on models with streaming: true. A streamed call makes one image.
  • Checked first. n, resolution, quality and the other listed fields are checked against the model before anything is reserved, so a value it doesn't take answers 400 for free.

Ten 4K images come to well over 100 MB of base64. Add "response_format": "url" and each image is stored and answered as a link instead, so the answer stays small:

{
  "created": 1790809771,
  "data": [
    {
      "url": "https://...r2.cloudflarestorage.com/binference-images/results/42/...png?X-Amz-...",
      "media_type": "image/png",
      "expires_at": "2026-10-08T12:00:00.000Z"
    }
  ],
  "usage": { "cost": 0.0084 }
}

Each link works for 7 days, then the image is deleted. Links are free, and an agent may hold 5 GB of them at once. They come with a whole answer, not with stream. If storing ever fails, you get the images as base64 instead, never nothing.

What it costs

Each image model lists its price lines: per image, per megapixel or per token, some for one tier (2k) or quality (low_1k). You pay what the generation cost plus the same fee as every call, and usage.cost shows it. A generation that fails costs nothing.

The call reserves its worst case: every image it asks for, at the dearest host's price, at the resolution and quality it names. Leave them out and it reserves the largest tier and the highest quality the model offers, so set them to hold less.

Through chat

A few chat models draw too, such as the Gemini image models and GPT-5 Image. Add modalities on Chat Completions and the images come back in message.images, as data: URLs. On Responses, they come back as image_generation_call items.

{
  "model": "google/gemini-3.1-flash-image",
  "modalities": ["image", "text"],
  "max_tokens": 4096,
  "messages": [{ "role": "user", "content": "Draw a small green cactus, flat icon." }]
}

They're priced per image token (pricing.image_output in GET /models), and max_tokens bounds the reserve as on any call.

Image tools

Image generation offered as a tool inside a chat, such as OpenAI's image_generation, is not served: the model decides how many images to make, so the cost has no limit to reserve. When a client only lists it, as Codex does, it's left out of the request. To make images, call POST /images.

On this page