Three to ten seats,one box
The seat count where ownership beats subscriptions, emphatically.
The maths
At three seats the payback turns. At five it is emphatic. One machine serves the team — with the honest caveat that which model you share matters: the coding model pools eight ways at reading speed, the deep-reasoning flagship is a one-lane specialist you switch to. We size this with you instead of pretending one number covers both.
One box, everyone on it
The measured team pattern: Laguna-S-2.1 serves 8 parallel streams at ~15 tok/s each (121.9 aggregate, 262K context) — reading speed for every seat at once. The flagship at our measured quant serves one stream; it is the specialist on tap, not the pool. Your team shares a machine, not an account.
The stack we set up for this
One box, the whole team on it:
- Open WebUI with per-person accounts — the ChatGPT your team already knows
- Laguna-S-2.1 for team coding: 121.9 tok/s aggregate at 8 streams — ~15 tok/s per seat, 262K context
- gpt-oss-120b for high-volume serving: 286 tok/s sustained aggregate at 256 streams (batch throughput, not per-seat speed)
- The flagship on tap for heavy reasoning, switchable in seconds
- n8n automations calling the box as a node
The meter you stop feeding
Working-hours agent load runs $12,000–19,000/year on frontier rates. One machine from €5,360 all-in (cheapest verified chassis; the DGX Spark reference build is €6,450) ends that line item — and adds the thing no subscription sells: nothing your team types leaves the building. Unsure? Pilot one desk box for 90 days before committing the team.