Skip to main content

The library

Every model that fits, byte-verified

Sizes are summed from the actual files on Hugging Face, speeds are measured on real hardware, and the list of what does NOT fit is published right below the list of what does. That is the whole method. Last verified: 5 August 2026.

Text & agents

The working set for reasoning, coding and agent loops — largest fitting quant shown with its exact size.

DeepSeek V4 Flash 0731

DeepSeek · 304B MoE · 13B active

The frontier open flagship — deep reasoning and agentic work, uncensored.

UD-IQ3_XXS · 104.2GB · DS4 Q2 87GB · EXL3 106.9GB

MIT

Hands-on on one Spark: 103GB resident, 18GB free, 131K context, 15.5–16.6 tok/s decode, 411 tok/s prefill. The DS4 engine (antirez's DwarfStar, MIT) fits it in 87GB by quantising only the routed experts — its Blackwell fork reports 776 tok/s sustained prefill and 59 tok/s aggregate across 12 concurrent agents on one box (project-reported, Aug 2026). New community route (0xSero's SparkInfer recipe, Aug 2026): a REAP-pruned EXL3 build (106.9GB, byte-verified) serving 262K context on ONE box at 34.3–48.9 tok/s decode across five published trials (median 38.1) with speculative decoding, and a 1,055 tok/s cold prefill at 252K tokens — repo-reported with committed evidence and a pinned Docker; the project itself labels its 35 tok/s floor an open gate, and we have not run it on our bench yet. 1M context architecturally.

Weights on Hugging Face

Laguna S 2.1

Poolside · US · 118B MoE · 8B active

Agentic coding at frontier-adjacent quality — the US-made permissive pick.

NVFP4 71GB · Q6_K_XL 107.1GB

OpenMDW-1.1

Poolside's own launch: “small enough to run on a single NVIDIA DGX Spark.” Hands-on: 31 tok/s with speculative decoding, 121.9 tok/s aggregate at 8 streams, 262K context. FP8 (131.3GB) does not fit — NVFP4 only.

Weights on Hugging Face

Qwen3.5 122B A10B

Alibaba · 125B MoE · 10B active

The long-context workhorse — 262K measured on one box with ~52GB free at NVFP4.

UD-Q6_K_XL 112.4GB · NVFP4 75.6GB

Apache-2.0

Weights on Hugging Face

Inkling-Small

Thinking Machines · US · 276B

A second US frontier lab in the library — reasoning-heavy work.

UD-IQ3_S 107.5GB

Apache-2.0

Weights on Hugging Face

MiniMax-M2.5

MiniMax · 230B MoE · 10B active

High-throughput generalist — ~26 tok/s measured, 128K context.

UD-Q3_K_XL 101.3GB

Community

Requires --no-mmap on unified memory.

Weights on Hugging Face

gpt-oss-120b

OpenAI · US · 120B MoE · 5.1B active

The concurrency king: 286 tok/s sustained aggregate at 256 streams.

Q8_0 63.4GB — half the box free

Apache-2.0

Natively MXFP4 — every quant lands 62.6–64.5GB; quantising does not shrink it.

Weights on Hugging Face

Ornith-1.0 35B

Ornith AI (ex-DeepReinforce) · 35B MoE

Full-precision inference with room to spare — zero quantisation loss.

BF16 70.2GB — no quantisation at all

MIT

Weights on Hugging Face

Ling-3.0-flash

inclusionAI (Ant) · 127B MoE

Ant's brand-new flash-tier MoE — frontier-adjacent, MIT, huge headroom on one box.

IQ4_XS 66.4GB · NVFP4 81.4GB

MIT

Released 2026-08-02; added to this library 2026-08-05. Sizes are summed file bytes from community GGUF/NVFP4 builds. The OFFICIAL FP8 build is 128.4GB — it does not fit one box, but fits a Y Tower 192 or a two-Spark cluster whole. Our own on-box measurement pending.

Weights on Hugging Face

Nemotron 3.5 Lightning

NVIDIA · US · 30B MoE · A3B · Mamba-2 hybrid

NVIDIA's agent workhorse for this exact machine — a million tokens of context on one box.

NVFP4 21.6GB · BF16 65.8GB — 1M ctx

OpenMDW

Released 2026-08-11; added 2026-08-12. Sizes are summed file bytes. NVIDIA's model card names 1× DGX Spark (GB10) as the single-GPU deployment target and ships a vLLM recipe running the full 1,048,576-token context with DSpark speculative decoding tuned for Spark. 'Best for long-running autonomous agents' is the vendor's framing. No throughput figure — unmeasured on our bench.

Weights on Hugging Face

Ling-3.0-tiny

inclusionAI (Ant) · 7.9B MoE · A1.3B

The always-on small sidekick — reasoning and agent chores at a fraction of the memory.

BF16 15.8GB — official FP8 and INT4 too

MIT

Released 2026-08-10; added 2026-08-12. BF16 size is summed file bytes. inclusionAI validated it on DGX Spark themselves: 100–105 tok/s in FP8 with ~8.3GiB peak memory at 8K context — vendor-reported numbers, unmeasured on our bench.

Weights on Hugging Face

Gemma 4 26B-A4B

Google · US · 26B MoE · 3.8B active

The fast daily driver — a purpose-built DGX Spark GGUF, 3,125 downloads/month.

Q4_K_M 16.8GB — Spark-specific build

Gemma licence

Weights on Hugging Face

LFM2.5-2.6B

Liquid AI · US · 2.6B dense hybrid · conv + GQA

The tool-calling sidekick — plans, calls tools and runs multi-step loops at near-zero cost while the flagship thinks.

~3GB class — dozens fit alongside the flagship

LFM open-weight — read before commercial use

Trained with agentic RL, 128K context, ~34T tokens. Vendor-reported: ToolSandbox 77.83, Multi-IF 80.07, IFStruct 85.49 — ahead of Gemma-4-E4B on all three (Liquid's own harness; labelled as such). Built for phones — on this box it is effectively free to keep resident next to the flagship for cheap tool-call turns.

Weights on Hugging Face

Muse Glimmer 30B

Meta Superintelligence Lab · US · 30B · perception encoder

Meta's agentic model for consumer hardware — multi-step reasoning with vision, built to run where you are.

BF16 GGUF 55.7GB — full precision, half the box free · Q8_0 29.6GB

Apache-2.0

Released 2026-08-09; added 2026-08-11. Meta Superintelligence Lab's own framing: distilled from Muse Spark, purpose-built for autonomous agentic tasks on consumer hardware (vendor-reported; unmeasured on our bench). Sizes are summed file bytes; unsloth GGUFs from day one.

Weights on Hugging Face

Maple-Preview

DeepGrove · US · 20B MoE · A1B · ternary weights

A reasoning sidekick that costs almost nothing to keep resident next to the flagship.

TQ2_0 6.3GB · TQ1_0 5.0GB — official GGUFs

MIT

Released 2026-08-04; added 2026-08-07 once DeepGrove shipped official llama.cpp ternary GGUFs. Vendor-reported: SOTA reasoning for its weight class, IMO-level problems, 218 tok/s on an M4 Mac mini, 131K context — all the vendor's own framing, labelled as such; unmeasured on our hardware.

Weights on Hugging Face

KAT-Coder-V2.5-Dev

Kwaipilot (Kuaishou) · 35B MoE · A3B (Qwen3.6 base)

A dedicated agentic-coding model that runs whole at full precision on one box.

BF16 69.3GB — full precision, no quantisation

Apache-2.0

Released 2026-07-23; added 2026-08-06. Built on the Qwen3.6-35B-A3B base with an agentic-coding post-training recipe (technical report published). 69.3GB is summed file bytes — BF16 fits with ~50GB headroom for context. “Dev” checkpoint; benchmark claims are vendor-reported and we have not measured it on-box yet.

Weights on Hugging Face

What does not fit — and we will not pretend otherwise

MoE 'active parameters' is a compute figure, not a memory figure: all experts must be resident. Anyone selling you these on a desktop box is confusing the two.

  • Kimi K3 (2.8T)594GB even at 1-bit — Not on one box, not on four.
  • GLM-5.2 (744B)~405GB Int4-Int8Mix — Measured on 4× Spark at 22–42 tok/s — we build that cluster.
  • K-EXAONE 2.0 (750B MoE · A37B)Q2_K 272.3GB (community GGUF) — Apache-2.0 — fits the 4-box cluster with headroom; no published cluster run yet.
  • Hunyuan Hy3 (295B)IQ1_M 181.2GB — Two boxes minimum.
  • LongCat 2.0 (1.6T)no GGUF exists — Datacentre class.
  • Qwen3.8-2.4T (2.4T MoE · A95B)UD-Q1_0 397.3GB (unsloth) — Went open 2026-08-08 — fits the 4-box cluster at 1-bit; revenue-conditioned licence, no cluster run published yet.

Image generation — ten models, all measured on a real Spark

Generation time for a standard image via the official DGX Spark ComfyUI playbook. Every one fits.

ModelFamilyGenerationMemory
SD 3.5 mediumStable Diffusion34s22GB
SD 3.5 largeStable Diffusion82s29GB
FLUX.2-klein 9BFLUX.295s66GB
FLUX.1-schnellFLUX.1110s72GB
FLUX.1-devFLUX.1111s95GB
FLUX.1-Kontext-devFLUX.1111s66GB
Z-Image-TurboZ-Image178s22GB
Z-ImageZ-Image179s22GB
Qwen-Image-2512Qwen212s63GB
FLUX.2-dev (4-bit)FLUX.2397s34GB

Video, voice & music

The rest of the media stack — licences read before anything went on this page.

LTX-2.3

Lightricks · 22B DiT

Video generation with synced audio in one pass, on your desk. No geographic restriction.

~33GB working set (FP8 DiT + NVFP4 encoder)

LTX-2 Community — free < $10M ARR

Weights on Hugging Face

Wan 2.2

Alibaba · MoE video

The cleanest video licence in the field — Alibaba claims no rights over outputs.

Fits comfortably

Apache-2.0

Weights on Hugging Face

Whisper large-v3

OpenAI · 1.55B

Transcribes 99 languages, fully offline.

~3GB

MIT

Weights on Hugging Face

Parakeet TDT 0.6B v3

NVIDIA · 600M

Beats Whisper on European languages (6.32% vs 7.44% WER) at 5–10× the speed.

~600MB

CC-BY-4.0

Weights on Hugging Face

Mage-VL

Microsoft · US · Compact VLM · codec-native

Streaming video understanding — watches a live feed and answers about it, on-box.

BF16 9.5GB — summed file bytes

Apache-2.0

Released 2026-07-25; added 2026-08-06. Microsoft's codec-native streaming multimodal model (arXiv paper + GitHub). Capability claims are vendor-reported; not yet measured on our hardware.

Weights on Hugging Face

LFM2.5-VL-3B

Liquid AI · US · 3B VLM

The vision sidekick — reads screenshots, documents and photos on-box at tiny cost.

BF16 6.2GB — summed file bytes

LFM open-weight — read before commercial use

Released 2026-08-11; added 2026-08-13. Vision sibling of the LFM2.5-2.6B tool-caller already in this library, same lfm1.0 licence; 17 languages in the model card. Capability claims are vendor-reported; unmeasured on our hardware.

Weights on Hugging Face

North Micro Vision Instruct

Cohere · CA · 2.4B VLM · native-resolution

Cohere's compact multilingual vision model — reads images in 11 languages, built to fine-tune.

BF16 5.0GB — summed file bytes

Apache-2.0

Released 2026-08-10; added 2026-08-13. Apache-2.0 with native-resolution image support; the vendor positions it as a foundation for task-specific fine-tuning. Unmeasured on our bench.

Weights on Hugging Face

Audio8 TTS Preview 0.6B

Audio8 · 0.6B

Zero-shot voice cloning TTS in 11 languages — Apache, so commercially safe.

2.6GB total repo

Apache-2.0

Released 2026-07-28; added 2026-08-06. “SOTA-class” is the vendor's own framing — labelled as such. Preview checkpoint.

Weights on Hugging Face

Chatterbox

Resemble AI · 0.5B

Clones a voice from ~10 seconds of audio — MIT licensed, commercially safe.

Small

MIT

65.3% blind-test preference over ElevenLabs is vendor-reported.

Weights on Hugging Face

Kokoro-82M

hexgrad (community) · 82M

Fast, commercially-safe narration. No cloning — sometimes that is the feature.

~300MB

Apache-2.0

Weights on Hugging Face

ACE-Step v1.5 XL Turbo

ACE Studio · DiT 9.97GB

Full music tracks, locally — pre-bundled in the DGX Spark ComfyUI distro.

~15GB with encoders

Open weights

Weights on Hugging Face

Unlimited-OCR

Baidu · 3.3B

Document OCR at scale — contracts, scans, handwriting — entirely on the box.

6.7GB BF16 — runs alongside anything

MIT

2.7M monthly downloads on Hugging Face; the workhorse for the document pipelines our regulated-industry buyers run. Added 2026-08-05.

Weights on Hugging Face

Notably absent: MiniMax-H3, this summer’s hot video release — its licence excludes the EU, UK and US by name, which is why it is in our research file and not on this page. Licences get read here.

What they make — same brief, different models

Every image and clip below was generated from the identical prompt (“A compact matte-black AI mini computer on a walnut desk, golden-hour window light…”) using each model's published open weights — the differences you see are the models.

Output generated by FLUX.1-dev
FLUX.1-dev95GB on the box · 111s per image measured on a Spark
Output generated by Qwen-Image-2512
Qwen-Image-251263GB · 212s per image measured
Output generated by SD 3.5 Medium
SD 3.5 Medium22GB · 34s per image — the fast option
LTX videoThe video headline — ~33GB working set, free commercial under $10M ARR
Wan 2.2Apache-2.0 — the cleanest video licence in the field
Generated voice
ChatterboxMIT voice — clones from ~10s of audio
Generated music
ACE-Step v1.5A 30-second instrumental, generated locally-runnable weights

Demo batch generated from each model’s published weights via hosted inference — identical weights to what we install on the box. On-box generation times shown are separate hands-on measurements from the table above. We re-shoot this gallery on our own hardware the day the bench unit lands.

Questions, answered straight

How much of the 128GB can models actually use?
Operators consistently report 119–121GB allocatable after the OS and driver reserve. Our fit calls use ~115GB as the safe working pool, which is why a 104GB quant with KV cache headroom counts as fitting and a 118GB one does not.
Why do you quote 3-bit quants — doesn't quality collapse?
DeepSeek V4 Flash is quantisation-aware-trained with its routed experts natively in MXFP4, so repacking is unusually gentle. Unsloth's published figures show ~96% top-1 token agreement at Q4-class sizes. Below Q4 the published record stops — we are measuring that curve ourselves and will publish it.
What is the single best model to start with?
For agentic and reasoning work, DeepSeek V4 Flash 0731 at UD-IQ3_XXS — the heaviest thing that runs on one box, verified hands-on at 15.5–16.6 tok/s with 131K context. For coding with a Western licence chain, Laguna S 2.1 at NVFP4. We install both.
Can I run video and image models alongside a text model?
Not simultaneously at the large end — the text flagship occupies ~103GB. The practical pattern is switching workloads (models load in tens of seconds) or dedicating a second box to media. The install sets both patterns up.
Get it pre-installed