Provenance
Who measured what, on which harness
A benchmark score without its harness and its measurer is a decoration. We publish the three families separately — canonical leaderboard, independent measurement, vendor-reported — because scores across them cannot be compared. We learned this by pulling one of our own claims.
Terminal-Bench 2.1 — the canonical leaderboard
tbench.ai, 17 entries, scaffold shown per row because it moves scores by 3+ points for the same model. No open-weight model we ship appears here yet — that absence is information, and we say it rather than filling the gap with vendor numbers.
| # | Model | Accuracy | Scaffold |
|---|---|---|---|
| 1 | Fable 5 | 83.8% | Claude Code |
| 2 | GPT-5.5 | 83.1% | Codex |
| 3 | Fable 5 | 80.4% | Terminus 2 |
| 4 | Grok 4.5 | 79.3% | Cursor CLI |
| 5 | Opus 4.8 | 78.9% | Claude Code |
| 6 | GPT-5.6 Terra | 78.4% | Codex |
| 7 | GPT-5.5 | 78.0% | Terminus 2 |
| 8 | Muse Spark 1.1 | 76.2% | mini-SWE-agent |
| 9 | GPT-5.6 Luna | 75.7% | Codex |
| 10 | Sonnet 5 | 74.6% | Claude Code |
| 11 | Gemini 3 Pro | 73.9% | Terminus 2 |
| 17 | GLM-5.1 | 58.7% | Claude Code |
Note rows 1 and 3: the same Fable 5 scores 83.8 on one scaffold and 80.4 on another. That 3.4-point spread is why cross-harness comparisons are meaningless — and why most benchmark marketing is.
Vendor-reported open-model figures — labelled as exactly that
Real numbers from real releases, run by the vendor on the vendor's harness. Useful signal, never comparable to the table above.
The engine race — the same box keeps getting faster
DeepSeek V4 Flash on one Spark: our measured llama.cpp baseline against antirez's DS4 engine and its Blackwell fork. Their numbers are project-reported until we re-run them on our bench unit — labelled accordingly.
| Engine | Footprint | Prefill | Decode | Serving | Provenance |
|---|---|---|---|---|---|
| llama.cpp · UD-IQ3_XXS | 103GB | 411 tok/s | 15.5–16.6 tok/s | 1 stream | measured by us |
| DS4 (antirez) · Q2 | 87GB | 344 tok/s @7k | 13.8 tok/s | — | repo-reported |
| DS4 Blackwell fork · Q2 | 87GB | 776 sustained · ~1,010 peak | 29.9 tok/s | 59 tok/s agg @12 | fork-reported |
The fine print that the promo cards skip: at 12 concurrent agents each one gets ~4.9 tok/s (29.9 single-stream falls to 15.7 / 11.7 / 7.2 / 4.9 at c=2/4/8/12), and DS4 quantises only the routed experts — its 87GB Q2 is not the same artefact as our 103GB UD-IQ3_XXS, so quality comparisons across the two are not free. The fork’s quality table (MMLU 79.5, HumanEval 89.0, 70/70 needle at 130k) is its own harness through its own serving path. MIT-licensed, install is one line — sources: antirez/ds4 and Entrpi/ds4-on-spark. We install and pin whichever engine best fits your workload, and re-measure on your unit before handover.
The claim we pulled from our own front page
Because the method only means something if it costs us too.
We briefly ran “DeepSeek 82.7 beats Fable 80.5” as a headline. Then we traced it: 82.7 was vendor-reported on an unreleased harness, 80.5 was an independent measurement on a different scaffold — and Fable’s canonical best is 83.8. Not a comparison, so we pulled it. What survives every check is the position we actually sell: parity-class agentic capability at prices roughly seventy times apart — and zero marginal cost on your own hardware.
Questions, answered straight
- Why isn't DeepSeek V4 Flash on the canonical leaderboard?
- Nobody has submitted it with a standard scaffold yet — the canonical board skews to closed models with funded eval teams. Its absence is not evidence either way, which is precisely why we won't substitute vendor numbers and call them comparable.
- So how should a buyer actually decide?
- On verified fit (does the model load on the hardware — byte-checked), measured local throughput (hands-on numbers, published), the price gap (vendor rate cards, public), and your own tasks. We will run your workload on request rather than argue from someone else's harness.
- Will you publish your own benchmark runs?
- Yes — same discipline, stated scaffold, on the exact hardware we sell, the day our bench unit lands. Including the small-quant quality curve nobody has published.