The money question
Stop paying per token
Every number on this page is a published vendor rate or a hands-on measurement — and the page ends with the cases where a subscription is genuinely the better deal, because sending the wrong buyers away is why the right ones trust us.
What the meter charges, per million tokens
Read from each vendor's own pricing page. Cached-input rates included — a competent buyer caches, so our comparisons use the range.
| Model | Input | Cached input | Output |
|---|---|---|---|
| Claude Opus 5 | $15 | $1.50 | $75 |
| Claude Fable 5 | $10 | $1 | $50 |
| GPT-5.6 | $8.75 | $0.88 | $70 |
| Gemini 3.x Pro | $7 | $1.75 + storage/hr | $42 |
| DeepSeek V4 Flash (their API) | $0.14 | $0.0028 | $0.28 |
Sources, last verified 5 August 2026: Anthropic, OpenAI, Google, DeepSeek, Mistral.
Yes — DeepSeek’s own API undercuts everything, including owning the box. It is also inference on someone else’s servers, under someone else’s laws, with your prompts in someone else’s logs. The machine is how you keep the open-model price and the privacy at the same time.
The money question
A month of this machine, priced on their meter
The meter charges for every token read and every token written. An agent loop does both, all day. Here is what one box actually processes in a month — and what that same work costs rented.
179M
tokens processed per month — 36M written, 144M read
$1,940–$3,230
the same work on Fable 5’s meter, monthly, with and without prompt caching — $23,277–$38,755 a year, forever
€5,360
the machine, once — then the meter reads €0 for the rest of its life
Derived from measured single-stream figures on one box running DeepSeek V4 Flash — 16 tok/s generation, 411 tok/s reading — in a 4:1 read-to-write agent loop, against Fable 5’s published $10 / $50 per 1M rates ($1/M cached reads; both ends of the range shown, never averaged). Concurrency depends on the model and engine: lighter models reach ~286 tok/s sustained aggregate on this hardware, and the flagship at our measured 103GB quant serves one stream. The DS4 engine’s 87GB build changes that — its fork reports 59 tok/s aggregate across 12 concurrent agents (project-reported, not yet our measurement) — we still quote our measured floor, not a reported ceiling. Electricity: ~240W flat out, about €1.70/day at €0.30/kWh.
Questions, answered straight
- When is a subscription honestly the better deal?
- Light, bursty chat with nothing sensitive in it. If nothing runs while you sleep and your prompts could be postcards, a $20-200/month subscription beats a €5,360 machine on pure cost. The answer flips the moment agents loop or the data is not yours to share.
- What does electricity actually cost?
- About 240W flat out — roughly €1.70 per day at €0.30/kWh, ~€50/month at full saturation. It appears in our calculator's fine print because a cost page that hides a cost defeats itself.
- Why do you show a price range instead of one savings number?
- Because prompt caching cuts the API side by roughly 40% for exactly the workloads this machine targets, and comparing against the uncached rate would overstate the saving. We show $1,940–$3,230/month for a saturated loop and never average the ends.
- Does clustering two boxes double the throughput?
- No — and we say so before you buy. Clustering pools memory (256GB) for model fidelity and headroom; fabric-bound serving adds capacity, not speed. The second box is for running bigger quants and serving more streams, not for making one stream faster.