Base spec
0.0 · 6Findings returned against a control the bench marked clean.
CRYSTALBENCH v1.0
CrystalBench measures whether a model can safely build and edit a living software spec. We run 9 tests through the real product and score the resulting database state, not the model's claims.
Leaderboard
Rows are model-effort runs. Turn on Compact reasoning efforts to reduce 32 rows to 12. Weighted and flat rankings disagree, with Fable 5 · low leading flat. The 0.53-point gap is within rerun variation, so treat the leader as a close result.
| 1 | Kimi K3defaultcurrent fixture | 82.0 | 80.1 | $2.27 | 100.0k |
| 2 | Fable 5lowcurrent fixture | 81.4 | 82.1 | $11.32 | 63.5k |
| 3 | Opus 5xhicurrent fixture | 80.7 | 79.8 | $8.64 | 188.7k |
| 4 | Opus 5medcurrent fixture | 78.9 | 77.8 | $6.94 | 103.3k |
| 5 | Fable 5xhicurrent fixture | 78.8 | 78.0 | $16.61 | 167.5k |
| 6 | Fable 5medcurrent fixture | 78.6 | 77.6 | $12.62 | 90.6k |
| 7 | Fable 5highprevious fixture | 78.5 | 78.2 | $13.67 | 120.9k |
| 8 | Opus 5lowcurrent fixture | 75.8 | 74.9 | $6.01 | 69.5k |
| 9 | Sonnet 5lowcurrent fixture | 75.3 | 74.9 | $4.06 | 84.9k |
| 10 | Sonnet 5medcurrent fixture | 73.7 | 72.6 | $5.74 | 171.7k |
| 11 | DeepSeek V4 Flash 0731lowcurrent fixture | 72.4 | 70.9 | $0.09 | 165.3k |
| 12 | Terrahighcurrent fixture | 72.1 | 72.4 | $2.04 | 29.7k |
| 13 | Terramedcurrent fixture | 71.4 | 71.7 | $1.97 | 25.4k |
| 14 | Sonnet 5xhicurrent fixture | 71.4 | 69.8 | $8.08 | 239.4k |
| 15 | Opus 5highprevious fixture | 70.8 | 71.6 | $7.99 | 159.9k |
| 16 | DeepSeek V4 Flash 0731maxcurrent fixture | 70.6 | 69.7 | $0.14 | 320.3k |
| 17 | Terralowcurrent fixture | 70.0 | 70.0 | $1.98 | 26.0k |
| 18 | Sonnet 5highprevious fixture | 69.8 | 69.7 | $5.63 | 193.6k |
| 19 | DeepSeek V4 Flash 0731highcurrent fixture | 69.0 | 67.7 | $0.13 | 303.9k |
| 20 | Solxhicurrent fixture | 67.8 | 66.4 | $6.62 | 79.1k |
| 21 | Lunahighcurrent fixture | 67.8 | 68.6 | $0.26 | 82.1k |
| 22 | Lunamedcurrent fixture | 67.4 | 70.0 | $0.19 | 30.3k |
| 23 | Grok 4.5highcurrent fixture | 66.9 | 65.7 | $0.97 | 48.8k |
| 24 | Deepseek 4 Prodefaultprevious fixture | 64.9 | 64.7 | $1.20 | 284.2k |
| 25 | Terraxhiprevious fixture | 63.9 | 66.9 | $2.98 | 91.4k |
| 26 | Solhighcurrent fixture | 63.5 | 62.6 | $5.86 | 56.1k |
| 27 | Lunalowcurrent fixture | 63.2 | 64.2 | $0.18 | 23.7k |
| 28 | Sollowcurrent fixture | 61.9 | 63.7 | $5.02 | 27.8k |
| 29 | Solmedcurrent fixture | 61.5 | 62.5 | $5.33 | 42.9k |
| 30 | Lunaxhiprevious fixture | 60.2 | 61.3 | $0.38 | 140.4k |
| 31 | Gem 3.5FLdefaultprevious fixture | 60.0 | 59.5 | $0.55 | 166.2k |
| 32 | Qwen 3.7Fdefaultprevious fixture | 46.4 | 47.7 | $0.10 | 337.9k |
Scores are percentages, 0–100, graded by code against seeded database state. ▫ marks a row whose eight non-restraint benches ran on the previous fixture: rank it, but do not read a single cell against a ▪ row. Costs come from three different bases — see the run economics below. Output tokens are observed completion tokens for that row's nine-bench run: a workload and verbosity signal, not a quality score.
Weighting
These weights reflect both product importance and how much each bench varies across models. Scaffold carries 3×; defect finding, live edits, schema, and stability carry 2×; restraint and grounded QA carry 1.5×; transcription and reference resolution carry 1×. Grounded QA and reference resolution stay lower because their scores barely move; weights run from 1× to 3× and sum to 16, with observed ranges and the flat mean shown beside the weighted score.
Scaffold roundtrip build a full spec from a brain-dump | ×3 | 53.1–99.3 | the widest and hardest task and the product's flagship job — and the only bench that genuinely ranks anything, with 27 distinct values across 32 configurations; construction coverage and output restraint are published separately |
Proposal gauntlet exact spec edits via chat proposals | ×2 | 12.5–87.5 | the highest-stakes surface — a wrong proposal mutates a user's live spec — and it still spreads models across 12.5–87.5 |
Inconsistency needle find planted spec defects | ×2 | 0.0–80.0 | genuine analytical work and the job people actually buy an analyzer for, but three of its six planted defects are found by everything and three by nothing — 15 of 32 land on the same 66.7, so it cannot carry top weight |
False-positive discipline stay silent on clean specs — re-measured on fixed fixture v3 | ×1.5 | 6.7–100.0 | the paired half of the needle: a defect-finder without restraint is unusable in production, and it spans the widest range on the board (6.7–100) |
Schema stress structured output across surfaces | ×2 | 15.0–84.8 | structured output is the gate every other capability passes through; a model that cannot hold the schema cannot deliver the reasoning behind it |
Enumerated data model transcribe an explicit field list | ×1 | 0.0–82.0 | transcription, not judgement — and ten configurations land on exactly 81.5% |
Trap references resolve near-duplicate entity names | ×1 | 25.0–100.0 | high-stakes in principle, but 28 of 32 scored exactly 75.0, so it does not get extra weight |
Grounded QA literal answers from the spec, no invention | ×1.5 | 75.0–100.0 | what users trust the product for daily, but it produces only three distinct values across the whole campaign — 26 of 32 scored 87.5 — so it is close to a constant and is weighted like one |
Stability same analysis five times, scored on consistency | ×2 | 0.0–80.0 | it grades variance rather than raw capability, but this campaign showed variance is the thing that actually bites — 23 distinct values across 32 configurations, and the only bench that repeats its own measurement |
The observed range is the lowest and highest score any configuration reached on that bench. A tenth bench, context scaling, was excluded for every model: it never produced a graded case, so it is not on this page and not in any score.
Per-bench breakdown
The matrix shows every configuration across all 9 benches, starting in leaderboard order. The second header row shows each bench's weight; the final columns show weighted and flat scores. Scaffold and false-positive discipline drive most ranking movement. Trap references and grounded QA stay nearly flat, but they remain because the behaviours matter even when they add little ranking information.
| ×3 | ×2 | ×2 | ×1.5 | ×2 | ×1 | ×1 | ×1.5 | ×2 | |||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Kimi K3defaultcurrent fixture | 99.3 | 75.0 | 80.0 | 60.0 | 84.8 | 81.5 | 75.0 | 87.5 | 78.0 | 82.0 | 80.1 |
| Fable 5lowcurrent fixture | 93.5 | 87.5 | 60.0 | 86.7 | 65.0 | 82.0 | 100.0 | 87.5 | 77.0 | 81.4 | 82.1 |
| Opus 5xhicurrent fixture | 98.5 | 75.0 | 66.7 | 73.3 | 68.2 | 81.5 | 75.0 | 100.0 | 80.0 | 80.7 | 79.8 |
| Opus 5medcurrent fixture | 99.0 | 75.0 | 66.7 | 60.0 | 63.6 | 81.5 | 75.0 | 100.0 | 79.0 | 78.9 | 77.8 |
| Fable 5xhicurrent fixture | 97.9 | 87.5 | 60.0 | 66.7 | 65.0 | 81.5 | 75.0 | 100.0 | 68.0 | 78.8 | 78.0 |
| Fable 5medcurrent fixture | 97.7 | 75.0 | 66.7 | 73.3 | 63.6 | 81.5 | 75.0 | 87.5 | 78.0 | 78.6 | 77.6 |
| Fable 5highprevious fixture | 91.7 | 87.5 | 60.0 | 86.7 | 62.5 | 81.5 | 75.0 | 87.5 | 71.6 | 78.5 | 78.2 |
| Opus 5lowcurrent fixture | 99.3 | 50.0 | 66.7 | 66.7 | 66.7 | 82.0 | 75.0 | 87.5 | 80.0 | 75.8 | 74.9 |
| Sonnet 5lowcurrent fixture | 93.9 | 50.0 | 66.7 | 60.0 | 66.7 | 82.0 | 75.0 | 100.0 | 80.0 | 75.3 | 74.9 |
| Sonnet 5medcurrent fixture | 93.4 | 62.5 | 66.7 | 40.0 | 66.7 | 81.5 | 75.0 | 87.5 | 80.0 | 73.7 | 72.6 |
| DeepSeek V4 Flash 0731lowcurrent fixture | 99.0 | 50.0 | 66.7 | 33.3 | 65.0 | 82.0 | 75.0 | 87.5 | 80.0 | 72.4 | 70.9 |
| Terrahighcurrent fixture | 84.1 | 50.0 | 60.0 | 80.0 | 62.5 | 74.1 | 75.0 | 87.5 | 78.0 | 72.1 | 72.4 |
| Terramedcurrent fixture | 87.3 | 50.0 | 60.0 | 80.0 | 65.0 | 76.8 | 75.0 | 87.5 | 63.7 | 71.4 | 71.7 |
| Sonnet 5xhicurrent fixture | 98.7 | 50.0 | 66.7 | 26.7 | 65.0 | 81.5 | 75.0 | 87.5 | 77.2 | 71.4 | 69.8 |
| Opus 5highprevious fixture | 73.0 | 75.0 | 66.7 | 46.7 | 60.0 | 81.5 | 75.0 | 100.0 | 66.7 | 70.8 | 71.6 |
| DeepSeek V4 Flash 0731maxcurrent fixture | 99.0 | 50.0 | 66.7 | 60.0 | 41.7 | 81.9 | 75.0 | 75.0 | 78.3 | 70.6 | 69.7 |
| Terralowcurrent fixture | 83.0 | 50.0 | 60.0 | 60.0 | 65.0 | 74.6 | 75.0 | 87.5 | 75.0 | 70.0 | 70.0 |
| Sonnet 5highprevious fixture | 87.0 | 62.5 | 54.5 | 46.7 | 62.5 | 81.5 | 75.0 | 87.5 | 69.8 | 69.8 | 69.7 |
| DeepSeek V4 Flash 0731highcurrent fixture | 96.2 | 37.5 | 66.7 | 20.0 | 65.0 | 82.0 | 75.0 | 87.5 | 79.0 | 69.0 | 67.7 |
| Solxhicurrent fixture | 84.9 | 62.5 | 66.7 | 6.7 | 63.6 | 72.8 | 75.0 | 87.5 | 78.0 | 67.8 | 66.4 |
| Lunahighcurrent fixture | 67.9 | 50.0 | 66.7 | 53.3 | 66.7 | 72.2 | 75.0 | 87.5 | 78.0 | 67.8 | 68.6 |
| Lunamedcurrent fixture | 53.1 | 62.5 | 60.0 | 80.0 | 66.7 | 76.8 | 75.0 | 87.5 | 68.6 | 67.4 | 70.0 |
| Grok 4.5highcurrent fixture | 91.7 | 50.0 | 60.0 | 6.7 | 63.6 | 82.0 | 75.0 | 87.5 | 74.9 | 66.9 | 65.7 |
| Deepseek 4 Prodefaultprevious fixture | 98.5 | 37.5 | 40.0 | 66.7 | 38.6 | 81.3 | 66.7 | 87.5 | 65.5 | 64.9 | 64.7 |
| Terraxhiprevious fixture | 56.7 | 50.0 | 40.0 | 86.7 | 58.3 | 76.8 | 75.0 | 87.5 | 71.1 | 63.9 | 66.9 |
| Solhighcurrent fixture | 86.0 | 50.0 | 66.7 | 6.7 | 63.6 | 75.5 | 75.0 | 87.5 | 52.7 | 63.5 | 62.6 |
| Lunalowcurrent fixture | 62.7 | 50.0 | 66.7 | 80.0 | 65.0 | 74.2 | 50.0 | 87.5 | 42.1 | 63.2 | 64.2 |
| Sollowcurrent fixture | 55.3 | 50.0 | 54.5 | 40.0 | 63.6 | 72.8 | 75.0 | 87.5 | 74.3 | 61.9 | 63.7 |
| Solmedcurrent fixture | 61.0 | 50.0 | 60.0 | 20.0 | 63.6 | 72.8 | 75.0 | 87.5 | 72.4 | 61.5 | 62.5 |
| Lunaxhiprevious fixture | 60.3 | 50.0 | 54.5 | 20.0 | 63.6 | 71.6 | 75.0 | 87.5 | 69.0 | 60.2 | 61.3 |
| Gem 3.5FLdefaultprevious fixture | 85.4 | 12.5 | 60.0 | 93.3 | 15.0 | 81.0 | 25.0 | 87.5 | 76.0 | 60.0 | 59.5 |
| Qwen 3.7Fdefaultprevious fixture | 53.9 | 62.5 | 0.0 | 100.0 | 50.0 | 0.0 | 75.0 | 87.5 | 0.0 | 46.4 | 47.7 |
Cell shading is the score itself, one hue, no thresholds. Bold is the column best. Every value is a single measurement: the same bench has swung more than thirty points between identical runs elsewhere in this campaign, so treat close cells as a tie. ▫ rows had their eight non-restraint benches measured on the previous fixture — the restraint column was re-measured on the current fixture for everyone, so that one column is comparable across the whole table.
Reasoning effort
7 models ran 3–4 reasoning efforts each. Going from a model's cheapest setting to its dearest improved the weighted score by at most 6.0 points, and 5 of 7 got worse, while runtime grew 2.0–4.1×. For 5 of 7, the most expensive setting is not the best one. A full rerun flipped one model's effort order too, so the effect is small and unstable; high and xhigh have not shown a repeatable benefit over medium here. Other workloads are untested.
| low current fixture | 72.4 | 21.0m | |
|---|---|---|---|
| high current fixture | 69.0 | 40.5m | |
| max current fixture | 70.6 | 41.3m |
−1.8 quality · 2.0× time
| low current fixture | 81.4 | 11.9m | |
|---|---|---|---|
| med current fixture | 78.6 | 17.6m | |
| high previous fixture | 78.5 | 22.9m | |
| xhi current fixture | 78.8 | 31.2m |
−2.6 quality · 2.6× time
| low current fixture | 75.8 | 13.4m | |
|---|---|---|---|
| med current fixture | 78.9 | 20.0m | |
| high previous fixture | 70.8 | 30.4m | |
| xhi current fixture | 80.7 | 35.2m |
+4.9 quality · 2.6× time
| low current fixture | 75.3 | 13.1m | |
|---|---|---|---|
| med current fixture | 73.7 | 25.1m | |
| high previous fixture | 69.8 | 30.6m | |
| xhi current fixture | 71.4 | 48.5m |
−4.0 quality · 3.7× time
| low current fixture | 63.2 | 11.3m | |
|---|---|---|---|
| med current fixture | 67.4 | 14.9m | |
| high current fixture | 67.8 | 30.2m | |
| xhi previous fixture | 60.2 | 46.5m |
−3.0 quality · 4.1× time
| low current fixture | 70.0 | 11.7m | |
|---|---|---|---|
| med current fixture | 71.4 | 13.9m | |
| high current fixture | 72.1 | 14.8m | |
| xhi previous fixture | 63.9 | 32.3m |
−6.1 quality · 2.8× time
| low current fixture | 61.9 | 14.3m | |
|---|---|---|---|
| med current fixture | 61.5 | 34.0m | |
| high current fixture | 63.5 | 24.9m | |
| xhi current fixture | 67.8 | 34.8m |
+6.0 quality · 2.4× time
Thick bar: weighted score on a 0–100 scale. Thin bar underneath: wall clock, scaled against the longest run on the board (48.5 minutes). The footer compares the lowest effort with the highest. Read only ▪-to-▪ pairs as a clean effort effect — a ▫ row had its eight non-restraint benches measured on the previous fixture. Wall clock is operational time, not model latency: the hosted model group is throttled and Anthropic/OpenAI runs pay process-spawn overhead per call.
Max output
An ideal router searches every team built from 31 eligible configurations and gives each bench to the team's strongest member. Qwen 3.7 Flash is excluded because its single 100% restraint result would skew the search; the next-best restraint configuration fills that slot. The best 3-model team reaches 89%, up from a solo baseline of 82%. That is a 7.1-point gain. The pale line marks solo performance; cyan shows what routing adds.
1 model
82%
2 models
87.6%+5.66
3 models
89%+7.08
The best three-model team combines:
Fable 5 · low earns its place as a complement to the other two, not as a standalone recommendation. This is a perfect-router ceiling: it assumes every task reaches the strongest member, and a different workload could produce a different trio.
| Scaffold roundtrip | Kimi K3defaultcurrent fixture | 99.3 | ×3 |
| Proposal gauntlet | Fable 5lowcurrent fixture | 87.5 | ×2 |
| Inconsistency needle | Kimi K3defaultcurrent fixture | 80.0 | ×2 |
| False-positive discipline | Fable 5lowcurrent fixture | 86.7 | ×1.5 |
| Schema stress | Kimi K3defaultcurrent fixture | 84.8 | ×2 |
| Enumerated data model | Fable 5lowcurrent fixture | 82.0 | ×1 |
| Trap references | Fable 5lowcurrent fixture | 100.0 | ×1 |
| Grounded QA | Opus 5xhicurrent fixture | 100.0 | ×1.5 |
| Stability | Opus 5xhicurrent fixture | 80.0 | ×2 |
This is a ceiling, not a result. It assumes a router that always picks correctly — nothing here demonstrates such a router, and building one is its own hard problem. Every cell behind it is a single measurement with real run-to-run variance. Treat 89% as the headroom a routing strategy could chase, not a score any deployment has achieved.
Cost vs quality
Each dot is one full suite. Cost runs horizontally on a log scale; weighted score runs vertically. Hover any dot to see which reasoning effort it was. The connected lines walk one model up its own effort ladder, and their flatness is how little extra score spending more buys. DeepSeek V4 Flash 0731 · low costs $0.09 for 72.4%; Fable 5 · xhigh costs $16.61 for 78.8%. The cheapest row is DeepSeek V4 Flash 0731 · low at $0.09, ranked 11th of 32.
Dot size is reasoning effort — smallest low, largest xhigh; the hosted provider models ran at their provider's default and appear once. Costs are shown as the dollar total for each suite. OpenAI's models actually bill to a subscription, so their dollars answer what the same work would have cost at published usage rates, not what was charged.
Run economics
One line per suite: every figure is summed from the per-call telemetry in that run's report, so a model measured at four settings appears four times, once for each. Read these as what a single nine-bench suite cost, not as a price list — the spread runs from about ten cents to seventeen dollars, and almost all of it is what each model charges per token rather than anything the benchmark did. Prompt tokens are the one figure the reports only aggregate per model, so this table shows completions.
| Qwen 3.7 Flash | default | 0.34M | $0.10 | 24.4m |
| DeepSeek V4 Flash 0731 | low | 0.17M | $0.09 | 21.0m |
| DeepSeek V4 Flash 0731 | high | 0.30M | $0.13 | 40.5m |
| DeepSeek V4 Flash 0731 | max | 0.32M | $0.14 | 41.3m |
| DeepSeek V4 Pro | default | 0.28M | $1.20 | 35.2m |
| Gemini 3.5 Flash-Lite | default | 0.17M | $0.55 | 8.1m |
| Grok 4.5 | high | 0.05M | $0.97 | 24.7m |
| Moonshot Kimi K3 | default | 0.10M | $2.27 | 97.6m |
| Claude Fable 5 | low | 0.06M | $11.32 | 11.9m |
| Claude Fable 5 | med | 0.09M | $12.62 | 17.6m |
| Claude Fable 5 | high | 0.12M | $13.67 | 22.9m |
| Claude Fable 5 | xhi | 0.17M | $16.61 | 31.2m |
| Claude Opus 5 | low | 0.07M | $6.01 | 13.4m |
| Claude Opus 5 | med | 0.10M | $6.94 | 20.0m |
| Claude Opus 5 | high | 0.16M | $7.99 | 30.4m |
| Claude Opus 5 | xhi | 0.19M | $8.64 | 35.2m |
| Claude Sonnet 5 | low | 0.08M | $4.06 | 13.1m |
| Claude Sonnet 5 | med | 0.17M | $5.74 | 25.1m |
| Claude Sonnet 5 | high | 0.19M | $5.63 | 30.6m |
| Claude Sonnet 5 | xhi | 0.24M | $8.08 | 48.5m |
| GPT-5.6 Luna | low | 0.02M | $0.18 | 11.3m |
| GPT-5.6 Luna | med | 0.03M | $0.19 | 14.9m |
| GPT-5.6 Luna | high | 0.08M | $0.26 | 30.2m |
| GPT-5.6 Luna | xhi | 0.14M | $0.38 | 46.5m |
| GPT-5.6 Terra | low | 0.03M | $1.98 | 11.7m |
| GPT-5.6 Terra | med | 0.03M | $1.97 | 13.9m |
| GPT-5.6 Terra | high | 0.03M | $2.04 | 14.8m |
| GPT-5.6 Terra | xhi | 0.09M | $2.98 | 32.3m |
| GPT-5.6 Sol | low | 0.03M | $5.02 | 14.3m |
| GPT-5.6 Sol | med | 0.04M | $5.33 | 34.0m |
| GPT-5.6 Sol | high | 0.06M | $5.86 | 24.9m |
| GPT-5.6 Sol | xhi | 0.08M | $6.62 | 34.8m |
| Campaign | 32 suites | 4.04M | $145.57 | 14.6h |
Wall clock is scored-bench time, an operational number rather than model latency: the hosted model group is throttled to 15 requests a minute, while Anthropic and OpenAI runs are unthrottled but pay process-spawn overhead on every call. One row to read with care — DeepSeek V4 Flash 0731 is the cheapest full suite at $0.09 and scores 72.4%; low cost is not the same as broad capability.
Why a separate benchmark
Public leaderboards measure whether a model can write a function. Ours asks whether it can change a living specification without quietly breaking it. The two skills barely know each other.
A flow step points at a role; a field points at its parent. Rename the wrong near-identical entity and nothing turns red — the damage is silent, and usually found weeks later by whoever tries to build from it.
A model that proposed nothing and a model that finished the job read exactly the same in a chat window. So we stopped taking anyone's word for it and started checking the database afterwards.
Nothing on this page was graded by another model. Each bench seeds its own workspace in an isolated database, runs against a fingerprinted fixture, and is scored by plain code — set overlap, precision and recall, or simply reading the project afterwards to see whether the change is there. A grader that cannot be charmed is a harsh one, which is why the scores are lower than you might expect.
Method and limits
A benchmark tells you about the shape of a problem long before it tells you anything reliable about a particular model. We would rather point at the thin spots ourselves than have you find them, so here they are, in the order that would embarrass us most.
The biggest surprise
Every other model here was measured once. Sol we measured twice — three complete suites, then three more a day later, same benches, same lane, nothing changed but the clock. Two of the three came back within about a point of where they started. The third moved by nearly seven, which was enough to change which setting looked best.
That is the closest thing to an error bar this page has, and it is why we would rather you read the ranking as a shape than a verdict. Nothing was wrong with either pass — this is simply what one measurement is worth.
| Bench | medium | high | xhigh | |||
|---|---|---|---|---|---|---|
| 1st → 2nd | Δ | 1st → 2nd | Δ | 1st → 2nd | Δ | |
| Scaffold roundtrip | 55.3 → 61.0 | +5.7 | 84.4 → 86.0 | +1.6 | 55.4 → 84.9 | +29.5 |
| Proposal gauntlet | 50.0 → 50.0 | — | 62.5 → 50.0 | −12.5 | 50.0 → 62.5 | +12.5 |
| Inconsistency needle | 54.5 → 60.0 | +5.5 | 60.0 → 66.7 | +6.7 | 61.5 → 66.7 | +5.2 |
| False-positive discipline | 40.0 → 20.0 | −20.0 | 6.7 → 6.7 | — | 40.0 → 6.7 | −33.3 |
| Schema stress | 62.5 → 63.6 | +1.1 | 63.6 → 63.6 | — | 37.5 → 63.6 | +26.1 |
| Enumerated data model | 72.3 → 72.8 | +0.5 | 72.3 → 75.5 | +3.2 | 72.3 → 72.8 | +0.5 |
| Trap references | 75.0 → 75.0 | — | 75.0 → 75.0 | — | 75.0 → 75.0 | — |
| Grounded QA | 87.5 → 87.5 | — | 87.5 → 87.5 | — | 87.5 → 87.5 | — |
| Stability | 74.9 → 72.4 | −2.5 | 76.4 → 52.7 | −23.7 | 69.4 → 78.0 | +8.6 |
Both passes are real provider runs through the product, graded by the same deterministic code (coverage-v1 + restraint-v1 for the scaffold column). The leaderboard shows the second pass. We publish the first rather than quietly replacing it, because the gap between them is the most useful number on this page.
Failure, in plain English
This is the archived false-positive-discipline run for Sol at high effort. The bench treats all three controls as clean, but Sol returned findings on every one. The score is the average of the three control scores: 0, 0, 20 → 6.7.
Findings returned against a control the bench marked clean.
Findings returned against a control the bench marked clean.
Findings returned against a control the bench marked clean.
This is a compact summary of the archived output, not a new model run. Sol mostly raised workflow consistency and handoff concerns; the full finding text stays in the report rather than on this page. Source: bench-parallel-2026-07-30T20-23-32-828Z.json.
Every score here is a single measurement, with no confidence interval behind it. So we tested what that is worth: we re-ran one model's three efforts end to end, all twenty-seven cells. Nine came back identical to the decimal. Four moved by more than twenty points. And the ranking of the three efforts inverted completely — the setting that had finished last finished first. The weighted totals held up far better than the cells inside them, which is the strongest argument on this page for reading the weighted column and ignoring any single bench. Treat a gap of a point or two as a tie.
Weights run 1× to 3× and are our editorial judgement about which benches represent real product work, not an output of the harness. They were revised once: grounded QA and inconsistency needle originally carried 3× each, until we counted how little they separate anything — three and six distinct values across 32 configurations — and moved that weight to scaffold and stability, which do. Re-weighting turned out to be nearly cosmetic, because a bench everyone ties on shifts everyone together; it changed no leader and moved nobody more than four places. We publish the flat unweighted mean next to the weighted one so you can see exactly what our judgement did to the ordering — and disagree with it.
The cost column shows a full-suite dollar total. Treat it as a rough comparison rather than an invoice, especially for OpenAI's subscription-billed models.
It runs against CrystalSpec's own AI surfaces and private fixtures, so it isn't something you can point at your own stack today. The methodology, though — that part you're very welcome to steal.
Every fixture behind CrystalBench is just an ordinary CrystalSpec project — typed flows, roles, data models, and test cases, all versioned and reviewable. Build one of your own and see what your coding agents make of it.
Start free