Built for founders and engineers who can't afford to have their stack subpoenaed.
We don't compete on the cheap end. If your monthly spend is under $100, Together, Hyperbolic, or DeepInfra will serve you better and faster. Our customer is the founding engineer at a stealth startup, the AI consultant under NDA, the team running production inference where the billing trail is its own threat model.
On May 13 2026, Anthropic raised Max-tier token caps and announced a $200/mo Agent SDK credit (effective 15 June) — because users like our operator were burning through the prior limits: 40+ million tokens in 17 of the past 21 days. That isn't abuse; that's what production inference looks like in 2026. The credit closes the raw-throughput gap for Max customers — but the structural friction remains. No US card on file. No prompts in a US discovery surface. No passport handed to a reseller. If any of that describes your situation, you're our customer.
We target production-grade developers on large and x-large projects: teams who have hit the ceiling on Anthropic, Cursor, Windsurf, and Cline, and who need an independent path on jurisdiction, payment rails, and data retention — not just on tokens-per-minute. We're the viable option for the workload they can't (or won't) run through the default stack — privacy, jurisdiction, and billing trail first; smart-route cost savings second.
Pay in crypto. Subscriptions renew monthly (cancel anytime). Credits never expire. We never read your prompts. We never charge a card.
FIDO2 required at sign-in (BYO) · why
Cancel anytime · 50% prorated refund within first 7 days of each month
Low-volume, single-use-case packages for hobbyists, solo devs, and small projects. BYO Keys add-on at $19/mo if you want to plug in your own provider keys.
opencode-go/deepseek-v4-flash — DeepSeek V4 chat at $0.14/$0.28 passthroughopencode-go/deepseek-v4-pro — DeepSeek V4-Pro reasoner at $0.435/$0.87 (75% reduction permanent 2026-05-22)opencode-go/gemini-3.5-flash — Day-0 in LiteLLM 1.86.2, $1.50/$9 passthrough~/.config/opencode/config.json at setup.html#opencodeStacks on any sub-$100 plan. Plug in your own Anthropic / OpenAI / Google / OpenRouter / Ollama keys — your wallet pays for inference on those routes, we do the routing.
Included free on all $100+ plans — Pro / Elite / Business / Scale / Sovereign / GLM-Dedicated / Stack already include BYO at no extra cost.
Narrower model menu, priced for one job. Pick a specialty if you know exactly what you're routing — Pro covers the general case at $119/mo.
Compliance-first routing + an optional frontier-escalation lane that stacks on any base subscription.
consortium-llama-70b, consortium-qwen-coder, consortium-qwen-235bPrefer to top up instead of subscribe? Same three tiers, same models, billed per 1M tokens with credits that never expire. Best for spiky workloads or first-time users who want to try before they commit monthly.
When subscription beats credits: if your monthly usage stays close to the included allowance (15M / 50M / 100M), subscription is cheaper per token at Pro and Elite. Credits are better for usage below 5M/mo or unpredictable bursts. Mix and match — same key, same models.
For teams at 250M+ tokens/month or operations needing reserved capacity and SLA commitments. Monthly billing in BTC.
Monthly subscription compared against routing the same workload to Sonnet 4.6 at retail.
| Monthly workload | All-Sonnet 4.6 | llmdeal tier | Savings |
|---|---|---|---|
| 15M tok/mo · personal coding assistant | $90/mo | Starter $59/mo | $31/mo · 34% |
| 50M tok/mo · light agent / prod assist | $300/mo | Pro $119/mo | $181/mo · 60% |
| 100M tok/mo · steady agentic work | $600/mo | Elite $159/mo | $441/mo · 74% |
| 300M tok/mo · team production | $1,800/mo | Business $499/mo | $1,301/mo · 72% |
| 1.5B tok/mo · heavy agent fleet | $9,000/mo | Scale $1,999/mo | $7,001/mo · 78% |
Subscription: each tier includes a monthly token allowance counted across input + output combined. Overage rolls onto the next 1M increment at the tier's overage rate.
Pay-as-you-go: top up any amount from $20; credits never expire. Each 1M tokens of usage is billed at the tier's per-1M rate (input + output counted separately).
Pro routing distributes traffic across ~50% Qwen-Coder, 25% Llama-3.3-70B, 15% DeepSeek-V3.2, 10% Codestral/Qwen3-235B/GLM-5 — we absorb model-to-model cost variance internally. Starter is single-model (our self-hosted Qwen-Coder-32B), no routing overhead.
Example A — Pro subscription at $119/mo includes 50M tokens.
Typical 800-in / 400-out request = 1,200 tokens.
50M / 1,200 = ~41,700 requests / month included.
Beyond 50M: $4 per additional 1M tokens.
Example B — Pro PAYG with $100 credit.
$2/M input + $5/M output, blended ~$3.50/M at 1:1.
$100 / $3.50 ≈ 28.5M tokens.
Credits never expire; top up anytime.Every tier routes across our self-hosted + open-weight stack (Llama, DeepSeek, Mistral, Qwen, GLM). Compared to running the same workload on Sonnet 4.6 retail: 34-78% lower cost depending on tier and volume — see the table above for the exact bucket.
Early builders who claimed a founder seat and pushed the product into shape before launch.
Software company shipping developer tools and AI workflow products — Snitch (security auditing), Jeremy (AI context layer), Scribe (voice-to-text), and a deep catalogue of focused developer and creative apps.
khuur.dev — precision software for developers and technical teams →
Real objections, straight answers — no sales spin.
The in-flight request completes — we absorb the overrun. Every subsequent request returns a 402 with an explicit "out of credits" body. Top up; service resumes immediately. No silent throttling, no surprise invoices.
Not today. Smart routing works because we hold the upstream contracts — that's what lets us route to the cheapest qualified model per request. BYO-key support is on the Sovereign tier roadmap, but it undercuts the routing margin, so it will be priced to reflect that.
A small open-source classifier (RouteLLM-style) scores each prompt on complexity, latency-sensitivity, code vs prose, and reasoning depth. Easy → Qwen-Coder-32B (our EU GPU). Fast workhorse → llama-3.3-70b-self-hosted (our EU GPU). Reasoning → DeepSeek V3.2 or Qwen3-Next 80B Thinking. Code-heavy → Codestral. Hardest queries → Qwen3 235B or GLM-5. Per-request telemetry shows exactly which model fired. The router is open-source and pinned in our repo — audit the logic yourself.
Elite defaults to EU-resident routing: every request is served from our EEA GPU and EU-resident open-weight model providers — when this routing is active, your prompts stay in the EEA and never enter US discovery scope. US capacity is available on request for Elite customers who prefer it. No frontier models (Claude, GPT, etc.) are provisioned on any tier — all models are self-hosted open-weight (Qwen, Llama, DeepSeek, etc.).
Refunds are paid in fiat (USD / EUR / SEK / NOK), not BTC. You receive the fiat value your crypto was worth on the day we recorded the inbound payment, minus per-second prorated usage. BTC price movement between purchase and refund is your exposure — we don't hedge FX. Refund window: cumulative usage < 3 hours across all orders ever (not calendar time). Fees shown in plain text before we send.
Pro+ accounts hold real spending power. A compromised account can drain credits faster than detection allows. FIDO2 (YubiKey, SoloKey, Apple/Google Passkeys, any FIDO2-compliant key) eliminates the phishing and credential-stuffing attack surface. Starter is FIDO2-optional; Pro / Elite require it at sign-in.
Yes, BYO. Bring your own FIDO2 key — we don't ship hardware. Most customers already have something that works:
Enrol your key from account settings on first sign-in. Details in privacy §8a.
When public preorder volume crosses $3,500, we fund a second EU GPU node — expanding capacity and adding larger open-weight models to the Pro routing pool. Pro has always routed exclusively across our self-hosted + open-weight stack (Llama, DeepSeek, Mistral, Qwen, GLM); the threshold unlocks more GPU headroom. Progress is tracked on the homepage public counter.
Every fiat rail (Visa, Stripe, ACH, SEPA, SWIFT) puts a KYC-bearing intermediary between you and us. The no-KYC promise becomes structurally unenforceable the moment a single fiat payment clears our books. Crypto keeps that chain broken.
Longer take: see the Why we don't accept fiat callout above.
Preorder $20 → $26 in credits and you get a working key the moment payment confirms (service is live 2026-05-22).
Pre-launch, the API base is https://llmdeal.me/v1 — run curl /v1/models
to see which models are live now. Model availability is public.