CRYSTALBENCH v1.0
16 models. 9 suites.
Know what each model is good at.
CrystalBench measures whether a model can safely build and edit a living software spec. We run 9 tests through the real product and score the resulting database state, not the model's claims.
- Overall winner
- Kimi K382% weighted
- Best under a dollar
- $0.43GLM 5.2 · default — 71.8%, 12th of 34
- Effort verdict
- UnsettledNo model peaks at medium; higher effort takes up to 4.1× as long, and one rerun flipped the result
- Routing ceiling
- 89%3 models routed per task, 7.1 above the best single model
Leaderboard
All model runs, ranked
- The rest
- AnthropicClaude models, one suite per effort
- OpenAIGPT models, one suite per effort
| 1 | Kimi K3defaultcurrent fixture | 82.0 | 80.1 | $2.27 | 100.0k |
| 2 | Fable 5lowcurrent fixture | 81.4 | 82.1 | $11.32 | 63.5k |
| 3 | Opus 5xhicurrent fixture | 80.7 | 79.8 | $8.64 | 188.7k |
| 4 | Sonnet 5lowcurrent fixture | 75.3 | 74.9 | $4.06 | 84.9k |
| 5 | Terrahighcurrent fixture | 72.1 | 72.4 | $2.04 | 29.7k |
| 6 | GLM 5.2defaultcurrent fixture | 71.8 | 68.9 | $0.43 | 79.3k |
| 7 | DeepSeek V4 Flash 0731defaultcurrent fixture | 70.7 | 69.4 | $0.12 | 263.2k |
| 8 | Solxhicurrent fixture | 67.8 | 66.4 | $6.62 | 79.1k |
| 9 | Lunahighcurrent fixture | 67.8 | 68.6 | $0.26 | 82.1k |
| 10 | Grok 4.5defaultcurrent fixture | 66.9 | 65.7 | $0.97 | 48.8k |
| 11 | Deepseek 4 Prodefaultprevious fixture | 64.9 | 64.7 | $1.20 | 284.2k |
| 12 | Minimax M3defaultcurrent fixture | 64.2 | 60.9 | $0.16 | 55.9k |
| 13 | Gemini 3.5 Flashdefaultprevious fixture | 60.0 | 59.5 | $0.55 | 166.2k |
| 14 | Tencent HY3defaultcurrent fixture | 59.8 | 59.0 | $0.22 | 306.3k |
| 15 | Qwen 3.8 Maxdefaultcurrent fixture | 59.6 | 60.6 | $2.69 | 322.9k |
| 16 | Qwen 3.7Fdefaultprevious fixture | 46.4 | 47.7 | $0.10 | 337.9k |
Scores are percentages, 0–100, graded by code against seeded database state. ▫ marks a row whose eight non-restraint benches ran on the previous fixture: rank it, but do not read a single cell against a ▪ row. Costs come from three different bases — see the run economics below. Output tokens are observed completion tokens for that row's nine-bench run: a workload and verbosity signal, not a quality score.
Weighting
Not every bench is worth the same
Scaffold roundtrip build a full spec from a brain-dump | ×3 | 53.1–99.3 | Build a complete project specification from a loose product brief. It checks whether the model captures the structure, details, and relationships the product needs. |
Proposal gauntlet exact spec edits via chat proposals | ×2 | 12.5–87.5 | Make precise changes to an existing specification through chat proposals. It checks whether the model edits the requested parts without damaging unrelated ones. |
Inconsistency needle find planted spec defects | ×2 | 0.0–80.0 | Find deliberate contradictions and missing links planted in a specification. It checks whether the model identifies real defects without inventing extra problems. |
False-positive discipline stay silent on clean specs — re-measured on fixed fixture v3 | ×1.5 | 0.0–100.0 | Read a clean specification and leave it alone when there is nothing wrong. It checks whether the model resists inventing defects or proposing unnecessary changes. |
Schema stress structured output across surfaces | ×2 | 15.0–84.8 | Return structured data in the shape the product expects across several specification surfaces. It checks whether the model keeps fields, types, and nesting intact while doing the work. |
Enumerated data model transcribe an explicit field list | ×1 | 0.0–82.0 | Turn an explicit list of fields into the product's data-model format. It checks faithful transcription of names, types, and optionality, with little room for interpretation. |
Trap references resolve near-duplicate entity names | ×1 | 25.0–100.0 | Resolve references when entity names are almost, but not exactly, the same. It checks whether the model links to the intended entity instead of taking a plausible near-match. |
Grounded QA literal answers from the spec, no invention | ×1.5 | 83.3–100.0 | Answer questions using only facts present in the specification. It checks whether the model finds the right detail and says when the answer is not there. |
Stability same analysis five times, scored on consistency | ×2 | 0.0–80.0 | Run the same analysis repeatedly on the same specification. It checks whether the result stays consistent instead of changing from one run to the next. |
The observed range is the lowest and highest score any configuration reached on that bench. A tenth bench, context scaling, was excluded for every model: it never produced a graded case, so it is not on this page and not in any score.
Per-bench breakdown
What drives the ranking
| ×3 | ×2 | ×2 | ×1.5 | ×2 | ×1 | ×1 | ×1.5 | ×2 | |||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Kimi K3defaultcurrent fixture | 99.3 | 75.0 | 80.0 | 60.0 | 84.8 | 81.5 | 75.0 | 87.5 | 78.0 | 82.0 | 80.1 |
| Fable 5lowcurrent fixture | 93.5 | 87.5 | 60.0 | 86.7 | 65.0 | 82.0 | 100.0 | 87.5 | 77.0 | 81.4 | 82.1 |
| Opus 5xhicurrent fixture | 98.5 | 75.0 | 66.7 | 73.3 | 68.2 | 81.5 | 75.0 | 100.0 | 80.0 | 80.7 | 79.8 |
| Opus 5medcurrent fixture | 99.0 | 75.0 | 66.7 | 60.0 | 63.6 | 81.5 | 75.0 | 100.0 | 79.0 | 78.9 | 77.8 |
| Fable 5xhicurrent fixture | 97.9 | 87.5 | 60.0 | 66.7 | 65.0 | 81.5 | 75.0 | 100.0 | 68.0 | 78.8 | 78.0 |
| Fable 5medcurrent fixture | 97.7 | 75.0 | 66.7 | 73.3 | 63.6 | 81.5 | 75.0 | 87.5 | 78.0 | 78.6 | 77.6 |
| Fable 5highprevious fixture | 91.7 | 87.5 | 60.0 | 86.7 | 62.5 | 81.5 | 75.0 | 87.5 | 71.6 | 78.5 | 78.2 |
| Opus 5lowcurrent fixture | 99.3 | 50.0 | 66.7 | 66.7 | 66.7 | 82.0 | 75.0 | 87.5 | 80.0 | 75.8 | 74.9 |
| Sonnet 5lowcurrent fixture | 93.9 | 50.0 | 66.7 | 60.0 | 66.7 | 82.0 | 75.0 | 100.0 | 80.0 | 75.3 | 74.9 |
| Sonnet 5medcurrent fixture | 93.4 | 62.5 | 66.7 | 40.0 | 66.7 | 81.5 | 75.0 | 87.5 | 80.0 | 73.7 | 72.6 |
| Terrahighcurrent fixture | 84.1 | 50.0 | 60.0 | 80.0 | 62.5 | 74.1 | 75.0 | 87.5 | 78.0 | 72.1 | 72.4 |
| GLM 5.2defaultcurrent fixture | 98.8 | 75.0 | 60.0 | 33.3 | 63.6 | 54.5 | 75.0 | 87.5 | 72.4 | 71.8 | 68.9 |
| Terramedcurrent fixture | 87.3 | 50.0 | 60.0 | 80.0 | 65.0 | 76.8 | 75.0 | 87.5 | 63.7 | 71.4 | 71.7 |
| Sonnet 5xhicurrent fixture | 98.7 | 50.0 | 66.7 | 26.7 | 65.0 | 81.5 | 75.0 | 87.5 | 77.2 | 71.4 | 69.8 |
| Opus 5highprevious fixture | 73.0 | 75.0 | 66.7 | 46.7 | 60.0 | 81.5 | 75.0 | 100.0 | 66.7 | 70.8 | 71.6 |
| DeepSeek V4 Flash 0731defaultcurrent fixture | 98.1 | 45.8 | 66.7 | 37.8 | 57.2 | 82.0 | 75.0 | 83.3 | 79.1 | 70.7 | 69.4 |
| Terralowcurrent fixture | 83.0 | 50.0 | 60.0 | 60.0 | 65.0 | 74.6 | 75.0 | 87.5 | 75.0 | 70.0 | 70.0 |
| Sonnet 5highprevious fixture | 87.0 | 62.5 | 54.5 | 46.7 | 62.5 | 81.5 | 75.0 | 87.5 | 69.8 | 69.8 | 69.7 |
| Solxhicurrent fixture | 84.9 | 62.5 | 66.7 | 6.7 | 63.6 | 72.8 | 75.0 | 87.5 | 78.0 | 67.8 | 66.4 |
| Lunahighcurrent fixture | 67.9 | 50.0 | 66.7 | 53.3 | 66.7 | 72.2 | 75.0 | 87.5 | 78.0 | 67.8 | 68.6 |
| Lunamedcurrent fixture | 53.1 | 62.5 | 60.0 | 80.0 | 66.7 | 76.8 | 75.0 | 87.5 | 68.6 | 67.4 | 70.0 |
| Grok 4.5defaultcurrent fixture | 91.7 | 50.0 | 60.0 | 6.7 | 63.6 | 82.0 | 75.0 | 87.5 | 74.9 | 66.9 | 65.7 |
| Deepseek 4 Prodefaultprevious fixture | 98.5 | 37.5 | 40.0 | 66.7 | 38.6 | 81.3 | 66.7 | 87.5 | 65.5 | 64.9 | 64.7 |
| Minimax M3defaultcurrent fixture | 95.4 | 62.5 | 50.0 | 0.0 | 61.1 | 46.7 | 75.0 | 87.5 | 70.1 | 64.2 | 60.9 |
| Terraxhiprevious fixture | 56.7 | 50.0 | 40.0 | 86.7 | 58.3 | 76.8 | 75.0 | 87.5 | 71.1 | 63.9 | 66.9 |
| Solhighcurrent fixture | 86.0 | 50.0 | 66.7 | 6.7 | 63.6 | 75.5 | 75.0 | 87.5 | 52.7 | 63.5 | 62.6 |
| Lunalowcurrent fixture | 62.7 | 50.0 | 66.7 | 80.0 | 65.0 | 74.2 | 50.0 | 87.5 | 42.1 | 63.2 | 64.2 |
| Sollowcurrent fixture | 55.3 | 50.0 | 54.5 | 40.0 | 63.6 | 72.8 | 75.0 | 87.5 | 74.3 | 61.9 | 63.7 |
| Solmedcurrent fixture | 61.0 | 50.0 | 60.0 | 20.0 | 63.6 | 72.8 | 75.0 | 87.5 | 72.4 | 61.5 | 62.5 |
| Lunaxhiprevious fixture | 60.3 | 50.0 | 54.5 | 20.0 | 63.6 | 71.6 | 75.0 | 87.5 | 69.0 | 60.2 | 61.3 |
| Gemini 3.5 Flashdefaultprevious fixture | 85.4 | 12.5 | 60.0 | 93.3 | 15.0 | 81.0 | 25.0 | 87.5 | 76.0 | 60.0 | 59.5 |
| Tencent HY3defaultcurrent fixture | 78.0 | 62.5 | 46.2 | 73.3 | 66.7 | 27.3 | 75.0 | 87.5 | 14.1 | 59.8 | 59.0 |
| Qwen 3.8 Maxdefaultcurrent fixture | 64.1 | 62.5 | 47.1 | 0.0 | 63.9 | 82.0 | 75.0 | 87.5 | 63.3 | 59.6 | 60.6 |
| Qwen 3.7Fdefaultprevious fixture | 53.9 | 62.5 | 0.0 | 100.0 | 50.0 | 0.0 | 75.0 | 87.5 | 0.0 | 46.4 | 47.7 |
Cell shading is the score itself, one hue, no thresholds. Bold is the column best. Every value is a single measurement: the same bench has swung more than thirty points between identical runs elsewhere in this campaign, so treat close cells as a tie. ▫ rows had their eight non-restraint benches measured on the previous fixture — the restraint column was re-measured on the current fixture for everyone, so that one column is comparable across the whole table.
Reasoning effort
More thinking costs time, not much score
Fable 5
| low current fixture | 81.4 | 11.9m | |
|---|---|---|---|
| med current fixture | 78.6 | 17.6m | |
| high previous fixture | 78.5 | 22.9m | |
| xhi current fixture | 78.8 | 31.2m |
−2.6 quality · 2.6× time
Opus 5
| low current fixture | 75.8 | 13.4m | |
|---|---|---|---|
| med current fixture | 78.9 | 20.0m | |
| high previous fixture | 70.8 | 30.4m | |
| xhi current fixture | 80.7 | 35.2m |
+4.9 quality · 2.6× time
Sonnet 5
| low current fixture | 75.3 | 13.1m | |
|---|---|---|---|
| med current fixture | 73.7 | 25.1m | |
| high previous fixture | 69.8 | 30.6m | |
| xhi current fixture | 71.4 | 48.5m |
−4.0 quality · 3.7× time
Luna
| low current fixture | 63.2 | 11.3m | |
|---|---|---|---|
| med current fixture | 67.4 | 14.9m | |
| high current fixture | 67.8 | 30.2m | |
| xhi previous fixture | 60.2 | 46.5m |
−3.0 quality · 4.1× time
Terra
| low current fixture | 70.0 | 11.7m | |
|---|---|---|---|
| med current fixture | 71.4 | 13.9m | |
| high current fixture | 72.1 | 14.8m | |
| xhi previous fixture | 63.9 | 32.3m |
−6.1 quality · 2.8× time
Sol
| low current fixture | 61.9 | 14.3m | |
|---|---|---|---|
| med current fixture | 61.5 | 34.0m | |
| high current fixture | 63.5 | 24.9m | |
| xhi current fixture | 67.8 | 34.8m |
+6.0 quality · 2.4× time
Thick bar: weighted score on a 0–100 scale. Thin bar underneath: wall clock, scaled against the longest run on the board (48.5 minutes). The footer compares the lowest effort with the highest. Read only ▪-to-▪ pairs as a clean effort effect — a ▫ row had its eight non-restraint benches measured on the previous fixture. Wall clock is operational time, not model latency: the hosted model group is throttled and Anthropic/OpenAI runs pay process-spawn overhead per call.
Max output
The best team mixes strengths
1 model
82%
- Kimi K3defaultcurrent fixture
2 models
87.6%+5.66
- Kimi K3defaultcurrent fixture
- Fable 5lowcurrent fixture
3 models
89%+7.08
- Kimi K3defaultcurrent fixture
- Fable 5lowcurrent fixture
- Opus 5xhicurrent fixture
Three models, one team
The best three-model team combines:
- Kimi K3defaultcurrent fixture
- Fable 5lowcurrent fixture
- Opus 5xhicurrent fixture
Fable 5 · low earns its place as a complement to the other two, not as a standalone recommendation. This is a perfect-router ceiling: it assumes every task reaches the strongest member, and a different workload could produce a different trio.
| Scaffold roundtrip | Kimi K3defaultcurrent fixture | 99.3 | ×3 |
| Proposal gauntlet | Fable 5lowcurrent fixture | 87.5 | ×2 |
| Inconsistency needle | Kimi K3defaultcurrent fixture | 80.0 | ×2 |
| False-positive discipline | Fable 5lowcurrent fixture | 86.7 | ×1.5 |
| Schema stress | Kimi K3defaultcurrent fixture | 84.8 | ×2 |
| Enumerated data model | Fable 5lowcurrent fixture | 82.0 | ×1 |
| Trap references | Fable 5lowcurrent fixture | 100.0 | ×1 |
| Grounded QA | Opus 5xhicurrent fixture | 100.0 | ×1.5 |
| Stability | Opus 5xhicurrent fixture | 80.0 | ×2 |
This is a ceiling, not a result. It assumes a router that always picks correctly — nothing here demonstrates such a router, and building one is its own hard problem. Every cell behind it is a single measurement with real run-to-run variance. Treat 89% as the headroom a routing strategy could chase, not a score any deployment has achieved.
Cost vs quality
39× the cost buys 7.0 more points
Each dot is one full suite. Cost runs horizontally on a log scale; weighted score runs vertically. Hover any dot to see which reasoning effort it was. The connected lines walk one model up its own effort ladder, and their flatness is how little extra score spending more buys. GLM 5.2 · default costs $0.43 for 71.8%; Fable 5 · xhigh costs $16.61 for 78.8%. The cheapest row is Qwen 3.7F · default at $0.10, ranked 34th of 34.
- The rest
- AnthropicClaude models, one suite per effort
- OpenAIGPT models, one suite per effort
Dot size is reasoning effort — smallest low, largest xhigh; the hosted provider models ran at their provider's default and appear once. Costs are shown as the dollar total for each suite, and they are not all the same kind of number: some are returned by the provider, the rest are measured tokens times a published rate. Where a group bills against a subscription rather than per token, its dollars answer what the same work would have cost at usage rates, not what was charged. The run economics table below breaks that down row by row.
Come build a spec worth benchmarking
Every fixture behind CrystalBench is just an ordinary CrystalSpec project — typed flows, roles, data models, and test cases, all versioned and reviewable. Build one of your own and see what your coding agents make of it.
Start free