vs RTX 5090
Four times faster. Four times less room.
The 5090 is a monster at what it can load — and it cannot load a single model in our library, at any quantisation anyone publishes. 32GB is the wall, and the interesting models live on the other side of it.
Buy the 5090 if your workload tops out at 35B
We mean it — and saying so is why you can trust the rest of this page.
Qwen3.5-35B at 165 tok/s against our 43. Gaming on the side. Image generation twice as fast. If 30–35B models cover your work, a 5090 workstation is the better machine and we will tell you so in the order flow. The Spark exists for the other case: DeepSeek V4 Flash at 104GB, Laguna at 107GB, MiniMax at 101GB — models a 5090 cannot open, at any quant, from any publisher.
Measured head-to-head — including where we lose
Same models, both machines, community-measured. The Spark takes the vision-language pair; the 5090 takes most of the rest of what fits it.
| Workload | RTX 5090 | DGX Spark | Winner |
|---|---|---|---|
| gpt-oss-20B | 1,338 tok/s | 1,094 tok/s | 5090 |
| Qwen3-4B class | 1,446 tok/s | 1,105 tok/s | 5090 |
| OCR 3B | 1,577 tok/s | 696 tok/s | 5090 |
| Qwen3-VL-4B | 1,005 tok/s | 1,237 tok/s | Spark |
| Qwen3-VL-8B | 868 tok/s | 972 tok/s | Spark |
| Qwen-Image (one image) | 46s | 98s | 5090 |
| ASR realtime factor | 0.324× | 0.342× | ≈ tie |
The rig maths
A single-5090 workstation lands around $7,500 — ~60% more than a Spark, ~4× the power draw, still walled at 32GB. Two 5090s with tensor split reach 64GB (enough for gpt-oss-120b) at roughly $12–15k, 1,150W of GPU draw, and tensor-parallel serving complexity. The 100GB class stays out of reach at any card count you would put under a desk.
The fit wall, concretely
Smallest published quants: DeepSeek V4 Flash 82.5GB · Laguna 54GB · MiniMax 93GB · gpt-oss-120b 59GB · Qwen3.5-122B 77GB. Against 32GB of VRAM, every one is a no — not slow, impossible. Memory is the axis where 2026’s interesting models moved, and cards did not move with it.
Questions, answered straight
- What about offloading to system RAM?
- CPU offload runs — at CPU speeds for the offloaded layers. Community measurements of big MoE models split across a 5090 and system RAM land in single-digit tok/s, below what the Spark does natively with the whole model in unified memory.
- Isn't the 5090 better value per token?
- For models that fit it, often yes. Value per token of a model that cannot load is zero — that is the whole comparison.