Skip to main content

Provenance

Who measured what, on which harness

A benchmark score without its harness and its measurer is a decoration. We publish the three families separately — canonical leaderboard, independent measurement, vendor-reported — because scores across them cannot be compared. We learned this by pulling one of our own claims.

Terminal-Bench 2.1 — the canonical leaderboard

tbench.ai, 17 entries, scaffold shown per row because it moves scores by 3+ points for the same model. No open-weight model we ship appears here yet — that absence is information, and we say it rather than filling the gap with vendor numbers.

#ModelAccuracyScaffold
1Fable 583.8%Claude Code
2GPT-5.583.1%Codex
3Fable 580.4%Terminus 2
4Grok 4.579.3%Cursor CLI
5Opus 4.878.9%Claude Code
6GPT-5.6 Terra78.4%Codex
7GPT-5.578.0%Terminus 2
8Muse Spark 1.176.2%mini-SWE-agent
9GPT-5.6 Luna75.7%Codex
10Sonnet 574.6%Claude Code
11Gemini 3 Pro73.9%Terminus 2
17GLM-5.158.7%Claude Code

Note rows 1 and 3: the same Fable 5 scores 83.8 on one scaffold and 80.4 on another. That 3.4-point spread is why cross-harness comparisons are meaningless — and why most benchmark marketing is.

Vendor-reported open-model figures — labelled as exactly that

Real numbers from real releases, run by the vendor on the vendor's harness. Useful signal, never comparable to the table above.

DeepSeek V4 Flash 073182.7DeepSeek's own run, unreleased harness
Laguna S 2.170.2Poolside's launch chart
DeepSeek V4 Flash (Apr preview)61.8matches Poolside's chart of the same variant — the two agree

The engine race — the same box keeps getting faster

DeepSeek V4 Flash on one Spark: our measured llama.cpp baseline against antirez's DS4 engine and its Blackwell fork. Their numbers are project-reported until we re-run them on our bench unit — labelled accordingly.

EngineFootprintPrefillDecodeServingProvenance
llama.cpp · UD-IQ3_XXS103GB411 tok/s15.5–16.6 tok/s1 streammeasured by us
DS4 (antirez) · Q287GB344 tok/s @7k13.8 tok/srepo-reported
DS4 Blackwell fork · Q287GB776 sustained · ~1,010 peak29.9 tok/s59 tok/s agg @12fork-reported

The fine print that the promo cards skip: at 12 concurrent agents each one gets ~4.9 tok/s (29.9 single-stream falls to 15.7 / 11.7 / 7.2 / 4.9 at c=2/4/8/12), and DS4 quantises only the routed experts — its 87GB Q2 is not the same artefact as our 103GB UD-IQ3_XXS, so quality comparisons across the two are not free. The fork’s quality table (MMLU 79.5, HumanEval 89.0, 70/70 needle at 130k) is its own harness through its own serving path. MIT-licensed, install is one line — sources: antirez/ds4 and Entrpi/ds4-on-spark. We install and pin whichever engine best fits your workload, and re-measure on your unit before handover.

The claim we pulled from our own front page

Because the method only means something if it costs us too.

We briefly ran “DeepSeek 82.7 beats Fable 80.5” as a headline. Then we traced it: 82.7 was vendor-reported on an unreleased harness, 80.5 was an independent measurement on a different scaffold — and Fable’s canonical best is 83.8. Not a comparison, so we pulled it. What survives every check is the position we actually sell: parity-class agentic capability at prices roughly seventy times apart — and zero marginal cost on your own hardware.

Questions, answered straight

Why isn't DeepSeek V4 Flash on the canonical leaderboard?
Nobody has submitted it with a standard scaffold yet — the canonical board skews to closed models with funded eval teams. Its absence is not evidence either way, which is precisely why we won't substitute vendor numbers and call them comparable.
So how should a buyer actually decide?
On verified fit (does the model load on the hardware — byte-checked), measured local throughput (hands-on numbers, published), the price gap (vendor rate cards, public), and your own tasks. We will run your workload on request rather than argue from someone else's harness.
Will you publish your own benchmark runs?
Yes — same discipline, stated scaffold, on the exact hardware we sell, the day our bench unit lands. Including the small-quant quality curve nobody has published.
See what verifiably runs