Private inference for onchain agents.
Every model runs inside a hardware enclave. Pay per call in stablecoins on Base, Robinhood Chain and Arc. No accounts.
Your prompts stay out of everyone else's logs.
Models run in an enclave
Inference happens inside an Intel TDX machine with a confidential NVIDIA GPU. The host operator cannot read prompts or outputs.
Every response is signed
Each call returns a receipt signed by the attested enclave, with a link to the public attestation report so you can check it yourself.
Nothing is cached or stored
We keep token counts for billing. We do not keep prompts, outputs or upstream error bodies.
Open models, served privately.
Prices are per million tokens, charged to the exact token. The list below is read live from the gateway.
Use it like any OpenAI endpoint.
Fund a key from your wallet
Sign in with a wallet, deposit from one dollar, and get a prepaid key. Close it any time and the balance is sent back.
Point your SDK at OffRouter
Change the base URL and the key. Streaming responses work the way your code already expects.
Or let the agent pay per call
With no key at all, the gateway answers 402 with a price. An x402 client signs it and the call goes through.
curl /v1/chat/completions \ -H "Authorization: Bearer $OFFROUTER_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek-v4-flash", "messages": [{"role": "user", "content": "Summarise this wallet history."}] }'
from openai import OpenAI
client = OpenAI(
base_url="/v1",
api_key=os.environ["OFFROUTER_KEY"],
)
reply = client.chat.completions.create(
model="kimi-k2.6",
messages=[{"role": "user", "content": "Summarise this wallet history."}],
)
print(reply.choices[0].message.content)
import { privateKeyToAccount } from "viem/accounts";
import { x402Client, wrapFetchWithPayment } from "@x402/fetch";
import { ExactEvmScheme } from "@x402/evm/exact/client";
const account = privateKeyToAccount(process.env.AGENT_KEY);
const client = x402Client.fromConfig({
// eip155:8453 Base, eip155:4663 Robinhood Chain, eip155:5042 Arc
schemes: [{ network: "eip155:4663", client: new ExactEvmScheme(account) }],
});
const pay = wrapFetchWithPayment(fetch, client);
const res = await pay("/v1/chat/completions", {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({
model: "gpt-oss-120b",
max_tokens: 512, // required: it bounds the amount you sign
messages: [{ role: "user", content: "Summarise this wallet history." }],
}),
});
Pay for what you use, where you hold funds.
The same price is quoted on every network. Your agent pays on whichever one it already has a balance on.
No account
A wallet is the only identity. There is no email, no card and no approval step between an agent and its first call.
A prepaid key is charged the real token count. The worst case is held first, so a key can never overrun its balance.
Pay per call signs a ceiling. Whatever the call did not use is sent back to the payer on the same network.
Holders pay less. Each tier takes a percentage off every call and adds a weekly allowance of free tokens.
See the tiers