Encrypted AI

Your agent's prompts and answers, encrypted on its own machine for a model running in sealed hardware. Nobody in between can read them, and every answer is signed by that hardware.

Sealed thinking

End-to-end encryptedSigned by the hardware

Your agent's words exist in only two places: on its own machine and inside the model's sealed hardware.

Who can read a prompt
  • Your agentReads it
  • bInferenceCiphertext only
  • NEAR AICan't look in
  • Cloud hostCan't look in
EncryptedOn your agent's machine
Runs inIntel TDX · NVIDIA Hopper GPUs
Every answerSigned in the hardware
Daily proofOn BNB Chain

With encrypted AI, your agent encrypts each message on its own machine for one model running inside a trusted execution environment (TEE): Intel TDX with NVIDIA confidential GPUs, on NEAR AI Cloud. bInference passes the call on and bills it, but sees only scrambled bytes, and so does everyone else on the way. The model's hardware decrypts it, answers, and encrypts the answer back for your agent alone.

Every answer is signed inside that hardware over the exact bytes your agent sent and received, so your agent can prove where it ran. It's for $BINF stakers: stake 500,000+ $BINF on the $BINF page, and each day your stake brings private AI credit. Encrypted calls spend only that credit, with your account's key.

What stays private

One call, end to end
Your agent

Encrypts every field for the model's attested key

Should my agent buy now?
Words
bInference

Checks your key and bills, passing the bytes on unchanged

9c1f04b7e2a58d3c…66b0e07a
Ciphertext
NEAR AI gateway

Routes the call to the model's hardware

9c1f04b7e2a58d3c…66b0e07a
Ciphertext
Model, sealed

Decrypts it inside, where nobody can look

Should my agent buy now?
Words
The prompt is encrypted before it leaves your agent. bInference and NEAR AI's gateway only ever hold ciphertext; the words appear again only inside the sealed hardware.

Tool definitions and tool calls are encrypted too, so a trading agent's tools and arguments never show. What can't be hidden: which model, when, how large, and which key paid.

Models

Four models NEAR AI runs in its own TEEs: GLM 5.3 Flash (z-ai/glm-5.3-flash), Qwen3.6 35B, Qwen3.8 27B and Qwen3 VL 30B. Each one's sealed machine was checked before we listed it, and a new model joins only once it is checked too. GET /api/private/v1/models lists them live, with their prices.

In your browser

The AI chat is the same private path for people: sign in with the wallet that stakes, pick one of the four models and ask. The page checks the model's hardware first, then encrypts each message before it leaves your browser. NVIDIA's service doesn't answer browsers, so the GPU evidence goes to it through bInference, and your browser checks NVIDIA's signature on the verdict itself.

  • No tools or web search. Nothing outside the sealed machine can read the chat, so nothing else can act on it.
  • Your chats stay in that browser. bInference keeps no copy and can't read one.
  • Each answer has a Privacy proof link, which opens its check at binference.io/proof.

Your account page shows your private AI apart from your other credit: what's left, what staking gave you in the last 7 days, what you used, and when the oldest part expires.

Use it

Install NEAR AI's official SDK, which verifies the hardware, encrypts and decrypts for you:

npm install @nearai/inference-sdk

Point it at bInference with your binf_ key:

import { InferenceClient } from "@nearai/inference-sdk/node";

const client = new InferenceClient({
  baseUrl: "https://binference.io/api/private/v1",
  apiKey: process.env.BINFERENCE_API_KEY,
  e2ee: true,
  signingAlgo: "ecdsa",
  gatewayVerification: { includeSpkiFingerprint: false },
});

const completion = await client.chat.completions.create({
  model: "z-ai/glm-5.3-flash",
  messages: [{ role: "user", content: "Should my agent buy now?" }],
});
console.log(completion.choices[0]?.message.content);

// Proof that this exact answer came from the attested hardware.
const proof = await client.verifyResponse(completion.id);
console.log(proof.signatureKind); // "provider_tee"
  • e2ee: true encrypts every message and tool field before it leaves your machine.
  • signingAlgo: "ecdsa" asks for signatures a contract can check on BNB Chain.
  • includeSpkiFingerprint: false: the TLS certificate your client sees is bInference's, not NEAR AI's. The hardware attestation is still verified in full.
  • Streaming works the same way (stream: true). Read the whole stream before verifyResponse.

What is checked

What is checked
  1. NEAR AI's gateway runs in sealed hardware

    Its Intel TDX quote, bound to a fresh nonce

    Your SDK
  2. The model runs in sealed hardware

    Its Intel TDX quote and its measured software

    Your SDKbInference
  3. Its GPUs run in confidential mode

    NVIDIA's own attestation service signs the verdict

    Your SDKbInference
  4. Every program in the machine is public code

    Each one since it started, matched to a signed public build on GitHub

    bInferenceAnyone
  5. The key you encrypt to was made inside

    Bound to the model's quote, so nobody else holds it

    Your SDK
  6. Every answer is signed over the exact bytes

    The SHA-256 of the request and of the answer, signed in the hardware

    Your SDKbInference
  7. Each day's answers and evidence on BNB Chain

    A Merkle root, and each key's full evidence, for anyone to check

    Anyone
Your agent's SDK checks the hardware before its first call and each answer after it. bInference also checks each signing key, and every program its machine ran, before it keeps an answer's receipt. Anyone can check an answer again at binference.io/proof.

After an answer, verifyResponse checks the signature over the exact bytes. bInference also checks every signing key itself the same way, and keeps a receipt of each signed answer: the model, the SHA-256 of the request and of the answer, and the signature. Never the prompt or the answer.

Every program in the machine. The SDK proves the hardware is genuine, but approves no particular software in it. bInference checks that too: every program the model's machine ran since it started is public code, either built and signed by GitHub from NEAR AI's public repos, or a public image pinned to one exact version. The machine's own log of what it started is hashed into an Intel-signed quote with our nonce, a fresh BNB Chain block, so the check shows it came after your answer.

Check any answer yourself at binference.io/proof: paste the answer's id, and your browser runs every check against Intel, NVIDIA, GitHub and BNB Chain, not against us.

Proof on BNB Chain. Once a day, the day's signed answers become one Merkle root posted to the ProofLog contract on BNB Chain, with each signing key's full evidence beside it. Anyone can check an answer against it, and the proof page does.

Errors

CodeHTTPWhat to do
encryption_required400This route takes encrypted calls only: use the SDK with e2ee: true
encryption_invalid400The encryption headers are incomplete: let the SDK send them
plaintext_fields400A message or tool field was not encrypted: the message names it
model_not_encrypted404The model is not served encrypted, or was named by an alias: use its exact id
stake_required403For $BINF stakers only: stake 500,000+ $BINF and use your account's key
private_unavailable503Encrypted models are briefly unavailable: retry later

Balance, key limits and rate limits answer as on every route: see Errors.

Good to know

  • Chat Completions only. The Responses and Messages formats can't be encrypted yet.
  • Hosted agents don't use encrypted AI yet: bInference writes their prompts, so it would see them.
  • Agent keys come later. For now the private models take your account's key and the AI chat, not the keys of agents you launched.
  • Plain calls to every model keep working at /api/v1, unchanged.
  • What a check can't see: the values of settings NEAR AI passes when it starts a program, such as where its metrics go. Their names are in the public config files, and the code that reads them is public.
  • The proof rests on Intel's and NVIDIA's hardware keys. Nothing that works without them is fast enough yet: fully encrypted computing is thousands of times too slow for chat models.

On this page