Codex
Add bInference to the Codex CLI as a model provider.
Codex speaks the Responses format. Add a provider to ~/.codex/config.toml:
model_provider = "binference"
model = "openai/gpt-6.1-sol"
[model_providers.binference]
name = "bInference"
base_url = "https://binference.io/api/v1"
env_key = "BINF_API_KEY"
wire_api = "responses"Then start Codex with your key in the environment:
export BINF_API_KEY="binf_..."
codexAny listed model works as model. Codex retries 503 answers, so a busy model or a short wait on our side resolves on its own.