Now live · Subscription or pay-as-you-go · BTC / XMR / LTC

The LLM API gateway your stack can't get subpoenaed.

US and EU GPU inference. Paid in BTC. No KYC. Pick a monthly subscription or top up pay-as-you-go credits — same key, same models, both stack. Choose EU-resident routing for EEA data residency, or use US capacity.

Subscription: Starter $59/mo · Pro $119/mo · Elite $159/mo · Business $499/mo · Scale $1,999/mo.
PAYG: $0.60/$1.20 per 1M (Starter) · $2/$5 (Pro) · $4/$9 (Elite). Credits never expire.

Who this is for

Built for founders and engineers who can't afford to have their stack subpoenaed.

We don't compete on the cheap end. If your monthly spend is under $100, Together, Hyperbolic, or DeepInfra will serve you better and faster. Our customer is the founding engineer at a stealth startup, the AI consultant under NDA, the team running production inference where the billing trail is its own threat model.

Why we exist

On May 13 2026, Anthropic raised Max-tier token caps and announced a $200/mo Agent SDK credit (effective 15 June) — because users like our operator were burning through the prior limits: 40+ million tokens in 17 of the past 21 days. That isn't abuse; that's what production inference looks like in 2026. The credit closes the raw-throughput gap for Max customers — but the structural friction remains. No US card on file. No prompts in a US discovery surface. No passport handed to a reseller. If any of that describes your situation, you're our customer.

We target production-grade developers on large and x-large projects: teams who have hit the ceiling on Anthropic, Cursor, Windsurf, and Cline, and who need an independent path on jurisdiction, payment rails, and data retention — not just on tokens-per-minute. We're the viable option for the workload they can't (or won't) run through the default stack — privacy, jurisdiction, and billing trail first; smart-route cost savings second.

  • Native BTC, XMR, LTC — zero KYC, on either side, ever.
  • EU or US GPU — your choice — opt into EU-resident routing and your prompts stay in the EEA, GDPR-compliant, never entering US discovery scope. Or run on US capacity. No geographic gate on who can sign up.
  • Smart routing across self-hosted + open-weight models — ~55–65% cheaper than all-Sonnet on the same workload.
  • Subscription, prepay, or pay-as-you-go — monthly plans from $15 to $1,999 with included tokens; 3-month and 12-month prepay packs for ~10-17% off (limited slots); or one-shot credit packs at the per-1M rate. Same key, same models, your call.

Pay in crypto. Subscriptions renew monthly (cancel anytime). Credits never expire. We never read your prompts. We never charge a card.

Starter

EU GPU inference · 15M tokens included · monthly
$59 / mo
monthly · 15M tokens included · 12-mo prepay $590 (5 slots)
or PAYG: $0.60 / $1.20 per 1M (see below)
  • Qwen2.5-Coder-32B on EEA GPU
  • OpenAI-compatible API — drop-in replacement, no SDK changes
  • 12k context (upgrading to 24-32k post-launch)
  • EU-resident inference available — enable it and your data stays in the EEA; US capacity also available
  • BTC / XMR / LTC checkout — no card, no KYC
  • Overage: $5 per additional 1M tokens (input or output)
  • Per-key budget caps — spending stays predictable
Subscribe Starter

Elite

70B + 128k context · 100M tokens included · monthly
$159 / mo
monthly · 100M tokens included · 12-mo prepay $1,590 (3 slots)
or PAYG: $4.00 / $9.00 per 1M (see below)
  • 128k context window
  • EU-resident routing by default — EEA GPU + EU-resident open-weight models; when this routing is active, your prompts stay in the EU. US capacity available on request.
  • GDPR Article 28 DPA available — sign before you send a single token
  • Dedicated rate limit pool — no contention with other tiers
  • Direct DM support (Matrix / Telegram) — reach a human, not a ticket queue
  • Extended EU model catalogue — additional 70B+ open-weight models available on request
  • Overage: $3 per additional 1M tokens
  • 🔐 FIDO2 required at sign-in — bring your own (YubiKey, SoloKey, Passkey).
  • All Pro features
Subscribe Elite

Cancel anytime · 50% prorated refund within first 7 days of each month

Entry-level bundles — under $100

Low-volume, single-use-case packages for hobbyists, solo devs, and small projects. BYO Keys add-on at $19/mo if you want to plug in your own provider keys.

$5 Trial Pack

$5 BTC minimum · ~5M tokens · never expires
$5 one-shot
~5M tokens · smart-route-fast lane
  • Minimum entry — $5 BTC top-up, no card, no KYC
  • ~5M tokens at Hobbyist rates (smart-route-fast / Llama 3.1 8B)
  • Credits never expire — top up more anytime
  • Real BTC payment verifies intent without locking you to a subscription
  • BYO add-on $19/mo if you need your own keys
Buy $5 Trial Pack

Hobbyist

2M / mo · Llama 3.1 8B · personal projects
$9 / mo
2M tokens · single model
  • Llama 3.1 8B on Groq (sub-second latency)
  • Best for: personal projects, learning, side-experiments
  • BTC / XMR / LTC checkout
  • BYO add-on $19/mo
Subscribe Hobbyist

Solo Dev

10M / mo · smart-route + fast
$23 / mo
10M tokens · smart-route + smart-route-fast · 12-mo prepay $230 (6 slots)
  • smart-route + smart-route-fast aliases
  • Best for: freelancers, solo devs, light agentic workloads
  • Sub-second latency on smart-route-fast
  • BYO add-on $19/mo
Subscribe Solo

Vision Pack

15M / mo · multimodal (image + text)
$47 / mo
15M tokens · Llama-4-Maverick (multimodal) + fast · 12-mo prepay $470 (4 slots)
  • NVIDIA Llama-4-Maverick — multimodal text + vision
  • smart-route-fast for text-only paths
  • Best for: designers, content tools, image-aware agents
  • Only sub-$100 tier with multimodal
  • BYO add-on $19/mo
Subscribe Vision

BYO Keys add-on · $19/mo

Stacks on any sub-$100 plan. Plug in your own Anthropic / OpenAI / Google / OpenRouter / Ollama keys — your wallet pays for inference on those routes, we do the routing.

Included free on all $100+ plans — Pro / Elite / Business / Scale / Sovereign / GLM-Dedicated / Stack already include BYO at no extra cost.

Specialty bundles

Narrower model menu, priced for one job. Pick a specialty if you know exactly what you're routing — Pro covers the general case at $119/mo.

Fast Pack

Low-latency · real-time chat · 50M / mo
$55 / mo
monthly · 50M tokens included · 12-mo prepay $550 (5 slots)
  • Llama 3.1 8B (Groq) — sub-second p50 latency
  • GPT-OSS-120B (Cerebras fast inference)
  • smart-route-fast alias
  • Best for: real-time chat, agent loops, IDE autocomplete
  • BTC / XMR / LTC checkout
Subscribe Fast

Coder Pack

Code-focused models · 30M / mo
$79 / mo
monthly · 30M tokens included · 12-mo prepay $790 (4 slots)
  • Qwen3-Coder-480B (256k context — large codebases)
  • Llama 3.3 70B (Groq + NVIDIA NIM)
  • smart-route-coder alias
  • Best for: devs, coding agents, IDE plugins (Continue / Cursor / Aider / Zed)
  • FIDO2 required at sign-in (BYO)
Subscribe Coder

Reasoner Pack

Reasoning + analysis · 25M / mo
$99 / mo
monthly · 25M tokens included · 12-mo prepay $990 (3 slots)
  • Qwen3-235B (Cerebras)
  • Nemotron-Super-120B (NVIDIA NIM)
  • GPT-OSS-120B (Cerebras)
  • smart-route-large alias
  • Best for: multi-step reasoning, research, long-form analysis
  • FIDO2 required at sign-in (BYO)
Subscribe Reasoner

EU-Sovereign and add-ons

Compliance-first routing + an optional frontier-escalation lane that stacks on any base subscription.

EU-Sovereign Pack

100% EU-resident · 60M / mo · DPA included
$199 / mo
monthly · 60M tokens included · 12-mo prepay $1,990 (3 slots)
  • EU-resident routing only — prompts never enter US discovery scope
  • GDPR Article 28 DPA signed before go-live
  • Sub-processor list documented + auditable
  • 90-day customer-readable audit log of every request
  • FIDO2 required at sign-in (BYO)
  • First customer signup triggers EU infrastructure provisioning
Subscribe EU-Sovereign

US-Sovereign Pack

100% US-resident · 60M / mo · US jurisdiction
$199 / mo
monthly · 60M tokens included · 12-mo prepay $1,990 (3 slots)
  • US-resident routing only — Groq + Cerebras + NVIDIA NIM US endpoints
  • US legal jurisdiction (Delaware shell entity available on request)
  • Sub-processor list documented + auditable
  • 90-day customer-readable audit log of every request
  • FIDO2 required at sign-in (BYO)
  • Compliance-friendly for US gov contractors / healthcare / fintech
  • Deliverable today — no new infra needed (most of consortium is US-hosted)
Subscribe US-Sovereign

Fleet

Small-team production · priority queue · 5 project keys
$519 / mo
max_parallel=20 · rpm=60 · tpm=200k · open-weight
  • All consortium open-weights + every smart-route alias
  • Priority queue — your calls jump the shared free-tier line
  • Higher concurrency: max_parallel=20 / rpm=60 / tpm=200k (sized to share upstream pool fairly)
  • 5 named project keys included (one per app/env)
  • BYO Keys included — Anthropic/OpenAI/Google/OpenRouter/Ollama
  • Custom router preferences (pin specific routes to specific upstreams)
  • 3-month prepay $1,399 (3 slots) · 12-month prepay $5,190 (2 founding slots, ~17% off)
Subscribe Fleet

Fleet Plus

Production + frontier credits · 60M frontier tokens / mo · live 2026-06-15
$959 / mo
max_parallel=40 · rpm=100 · tpm=400k · open-weight today, frontier 2026-06-15
  • Everything in Fleet
  • 60M frontier credits/month (Claude Sonnet 4.6 / GPT-5.5 / Gemini Pro)
  • Frontier escalation lane activates 2026-06-15 with the Anthropic Agent SDK credit lane; open-weight stack works today
  • Higher concurrency: max_parallel=40 / rpm=100 / tpm=400k
  • 10 named project keys included
  • Cache-priority — your repeated prompts hit the cache layer first
  • Custom router weights — bias smart-route to your preferred providers
  • 24h-priority Matrix/Discord support DM channel
  • 12-month prepay $9,590 (2 founding slots, ~17% off)
Subscribe Fleet Plus

Burst Day

24h at Fleet limits · 8M tokens any model · one-off
$129 one-time
single purchase · no subscription · no auto-renew
  • 24-hour window of Fleet-tier concurrency + priority queue
  • 8M tokens of any model — frontier included
  • One-off purchase, no subscription, no auto-renew
  • Perfect for a Sunday batch, a benchmark sweep, or a one-day codebase pass
  • Clock starts on your first API call after redemption
  • Stacks with any active subscription (limits add)
Get a Burst Day

Or pay as you go — per-1M-token credits

Prefer to top up instead of subscribe? Same three tiers, same models, billed per 1M tokens with credits that never expire. Best for spiky workloads or first-time users who want to try before they commit monthly.

Starter · PAYG

Qwen2.5-Coder-32B · single model
$0.60 / $1.20
per 1M tokens (input / output)
  • Top up in any amount from $20
  • Credits never expire
  • OpenAI-compatible API · same endpoint as subscription
  • BTC / XMR / LTC checkout · no KYC
  • Switch to subscription anytime
Buy Starter credits

Pro · PAYG

Smart-routed across 6 models
$2.00 / $5.00
per 1M tokens (weighted avg)
  • Top up in any amount from $50
  • Credits never expire
  • 32k context window · smart routing
  • Per-request cost telemetry
  • FIDO2 required at sign-in (BYO)
Buy Pro credits

Elite · PAYG

70B + 128k context · extended catalogue
$4.00 / $9.00
per 1M tokens (weighted avg)
  • Top up in any amount from $100
  • Credits never expire
  • 128k context · EU-resident routing default
  • Direct DM support (Matrix / Telegram)
  • FIDO2 required at sign-in (BYO)
Buy Elite credits

When subscription beats credits: if your monthly usage stays close to the included allowance (15M / 50M / 100M), subscription is cheaper per token at Pro and Elite. Credits are better for usage below 5M/mo or unpredictable bursts. Mix and match — same key, same models.

Business & Enterprise plans

For teams at 250M+ tokens/month or operations needing reserved capacity and SLA commitments. Monthly billing in BTC.

Business

$499 / mo
  • 300M tokens included monthly
  • Overage at $2 per additional 1M tokens
  • Reserved capacity during peak hours — no queue, guaranteed throughput
  • Priority Discord DM support
  • Monthly usage report broken down by model
  • Cancel anytime — no annual lock-in
  • Annual prepay option: $4,989/yr (1 month free)
Subscribe Business

Scale

$1,999 / mo
  • 1.5B tokens included monthly
  • Overage at $1.50 per additional 1M tokens
  • Dedicated Matrix/Slack channel with founder — direct line, no queue
  • Reserved capacity + 99.5% SLA
  • Weekly usage + cost breakdown by model
  • Quarterly architecture review
  • Annual prepay: $19,990/yr (2 months free)
Talk to founder

Sovereign

custom · annual
  • Dedicated routing pool — your own GPU slice on our EU or US box, your choice
  • GDPR Article 28 DPA + sub-processor list signed before go-live
  • Annual commit, invoiced — no card required
  • Direct operator phone / Signal line
  • Custom retention + audit terms
  • From $4,999 / year, scoped to your workload
Request quote

The cost case, plainly stated

Monthly subscription compared against routing the same workload to Sonnet 4.6 at retail.

Monthly workloadAll-Sonnet 4.6llmdeal tierSavings
15M tok/mo · personal coding assistant$90/moStarter $59/mo$31/mo · 34%
50M tok/mo · light agent / prod assist$300/moPro $119/mo$181/mo · 60%
100M tok/mo · steady agentic work$600/moElite $159/mo$441/mo · 74%
300M tok/mo · team production$1,800/moBusiness $499/mo$1,301/mo · 72%
1.5B tok/mo · heavy agent fleet$9,000/moScale $1,999/mo$7,001/mo · 78%
How the pricing works →

Subscription: each tier includes a monthly token allowance counted across input + output combined. Overage rolls onto the next 1M increment at the tier's overage rate.

Pay-as-you-go: top up any amount from $20; credits never expire. Each 1M tokens of usage is billed at the tier's per-1M rate (input + output counted separately).

Pro routing distributes traffic across ~50% Qwen-Coder, 25% Llama-3.3-70B, 15% DeepSeek-V3.2, 10% Codestral/Qwen3-235B/GLM-5 — we absorb model-to-model cost variance internally. Starter is single-model (our self-hosted Qwen-Coder-32B), no routing overhead.

Example A — Pro subscription at $119/mo includes 50M tokens.
  Typical 800-in / 400-out request = 1,200 tokens.
  50M / 1,200 = ~41,700 requests / month included.
  Beyond 50M: $4 per additional 1M tokens.

Example B — Pro PAYG with $100 credit.
  $2/M input + $5/M output, blended ~$3.50/M at 1:1.
  $100 / $3.50 ≈ 28.5M tokens.
  Credits never expire; top up anytime.

Every tier routes across our self-hosted + open-weight stack (Llama, DeepSeek, Mistral, Qwen, GLM). Compared to running the same workload on Sonnet 4.6 retail: 34-78% lower cost depending on tier and volume — see the table above for the exact bucket.

Pre-purchase FAQ

Real objections, straight answers — no sales spin.

What happens if credits hit zero mid-request?

The in-flight request completes — we absorb the overrun. Every subsequent request returns a 402 with an explicit "out of credits" body. Top up; service resumes immediately. No silent throttling, no surprise invoices.

Can I bring my own Anthropic / OpenAI key?

Not today. Smart routing works because we hold the upstream contracts — that's what lets us route to the cheapest qualified model per request. BYO-key support is on the Sovereign tier roadmap, but it undercuts the routing margin, so it will be priced to reflect that.

How does Pro routing decide which model fires?

A small open-source classifier (RouteLLM-style) scores each prompt on complexity, latency-sensitivity, code vs prose, and reasoning depth. Easy → Qwen-Coder-32B (our EU GPU). Fast workhorse → llama-3.3-70b-self-hosted (our EU GPU). Reasoning → DeepSeek V3.2 or Qwen3-Next 80B Thinking. Code-heavy → Codestral. Hardest queries → Qwen3 235B or GLM-5. Per-request telemetry shows exactly which model fired. The router is open-source and pinned in our repo — audit the logic yourself.

What does "EU-resident routing by default" actually mean on Elite?

Elite defaults to EU-resident routing: every request is served from our EEA GPU and EU-resident open-weight model providers — when this routing is active, your prompts stay in the EEA and never enter US discovery scope. US capacity is available on request for Elite customers who prefer it. No frontier models (Claude, GPT, etc.) are provisioned on any tier — all models are self-hosted open-weight (Qwen, Llama, DeepSeek, etc.).

How does the refund actually work in BTC?

Refunds are paid in fiat (USD / EUR / SEK / NOK), not BTC. You receive the fiat value your crypto was worth on the day we recorded the inbound payment, minus per-second prorated usage. BTC price movement between purchase and refund is your exposure — we don't hedge FX. Refund window: cumulative usage < 3 hours across all orders ever (not calendar time). Fees shown in plain text before we send.

Why is a FIDO2 hardware key required on Pro+? Do I bring my own?

Pro+ accounts hold real spending power. A compromised account can drain credits faster than detection allows. FIDO2 (YubiKey, SoloKey, Apple/Google Passkeys, any FIDO2-compliant key) eliminates the phishing and credential-stuffing attack surface. Starter is FIDO2-optional; Pro / Elite require it at sign-in.

Yes, BYO. Bring your own FIDO2 key — we don't ship hardware. Most customers already have something that works:

  • YubiKey 5 series — $45-75 from yubico.com or a local reseller (the gold standard)
  • SoloKey 2 / Token2 / NitroKey — $30-50, open-hardware alternatives
  • Apple / Google Passkey — free, if you have an iPhone / iPad / Mac with Touch/Face ID, or a Pixel / modern Android
  • Windows Hello + TPM — free on most modern Windows laptops

Enrol your key from account settings on first sign-in. Details in privacy §8a.

What's the $3,500 preorder threshold?

When public preorder volume crosses $3,500, we fund a second EU GPU node — expanding capacity and adding larger open-weight models to the Pro routing pool. Pro has always routed exclusively across our self-hosted + open-weight stack (Llama, DeepSeek, Mistral, Qwen, GLM); the threshold unlocks more GPU headroom. Progress is tracked on the homepage public counter.

Why no fiat — no cards, Stripe, or bank transfer?

Every fiat rail (Visa, Stripe, ACH, SEPA, SWIFT) puts a KYC-bearing intermediary between you and us. The no-KYC promise becomes structurally unenforceable the moment a single fiat payment clears our books. Crypto keeps that chain broken.

Longer take: see the Why we don't accept fiat callout above.

Can I test the API before paying?

Preorder $20 → $26 in credits and you get a working key the moment payment confirms (service is live 2026-05-22). Pre-launch, the API base is https://llmdeal.me/v1 — run curl /v1/models to see which models are live now. Model availability is public.

Available on demand