OpenAI SDK

The official OpenAI libraries for TypeScript and Python, with any model.

The OpenAI SDKs work with every model here, not only OpenAI's. Set baseURL and the key.

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://binference.io/api/v1",
  apiKey: process.env.BINF_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "google/gemini-3.8-flash",
  messages: [{ role: "user", content: "Write a haiku about BNB Chain." }],
  max_tokens: 512,
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Responses

client.responses.create works the same way, stateless: nothing is stored between calls, so send the conversation each time.

const reply = await client.responses.create({
  model: "openai/gpt-6.1-sol",
  input: "Summarize the last 24 hours of trades in one line.",
  max_output_tokens: 256,
});

console.log(reply.output_text);

Retries

The SDKs retry 429 and 5xx answers on their own, twice by default, and wait as long as the Retry-After-Ms header says. That is the right behavior here. Raise maxRetries (max_retries in Python) for jobs that run unattended. See Retries.

Other OpenAI-compatible tools

Anything that lets you set an OpenAI base URL works: LangChain's ChatOpenAI, LiteLLM, LlamaIndex and most agent frameworks. Use https://binference.io/api/v1 as the base URL and your binf_ key as the API key.

On this page