Skip to content

12 models. 9 suites.
Know what each model is good at.

CrystalBench measures whether a model can safely build and edit a living software spec. We run 9 tests through the real product and score the resulting database state, not the model's claims.

Overall winner
Kimi K382% weighted
Best under a dollar
$0.09DeepSeek V4 Flash 0731 · low — 72.4%, 11th of 32
Effort verdict
UnsettledNo model peaks at medium; higher effort takes up to 4.1× as long, and one rerun flipped the result
Routing ceiling
89%3 models routed per task, 7.1 above the best single model

All model runs, ranked

Rows are model-effort runs. Turn on Compact reasoning efforts to reduce 32 rows to 12. Weighted and flat rankings disagree, with Fable 5 · low leading flat. The 0.53-point gap is within rerun variation, so treat the leader as a close result.

  • The rest
  • Anthropic
  • OpenAI
1Kimi K3defaultcurrent fixture
82.0
80.1$2.27100.0k
2Fable 5lowcurrent fixture
81.4
82.1$11.3263.5k
3Opus 5xhicurrent fixture
80.7
79.8$8.64188.7k
4Opus 5medcurrent fixture
78.9
77.8$6.94103.3k
5Fable 5xhicurrent fixture
78.8
78.0$16.61167.5k
6Fable 5medcurrent fixture
78.6
77.6$12.6290.6k
7Fable 5highprevious fixture
78.5
78.2$13.67120.9k
8Opus 5lowcurrent fixture
75.8
74.9$6.0169.5k
9Sonnet 5lowcurrent fixture
75.3
74.9$4.0684.9k
10Sonnet 5medcurrent fixture
73.7
72.6$5.74171.7k
11DeepSeek V4 Flash 0731lowcurrent fixture
72.4
70.9$0.09165.3k
12Terrahighcurrent fixture
72.1
72.4$2.0429.7k
13Terramedcurrent fixture
71.4
71.7$1.9725.4k
14Sonnet 5xhicurrent fixture
71.4
69.8$8.08239.4k
15Opus 5highprevious fixture
70.8
71.6$7.99159.9k
16DeepSeek V4 Flash 0731maxcurrent fixture
70.6
69.7$0.14320.3k
17Terralowcurrent fixture
70.0
70.0$1.9826.0k
18Sonnet 5highprevious fixture
69.8
69.7$5.63193.6k
19DeepSeek V4 Flash 0731highcurrent fixture
69.0
67.7$0.13303.9k
20Solxhicurrent fixture
67.8
66.4$6.6279.1k
21Lunahighcurrent fixture
67.8
68.6$0.2682.1k
22Lunamedcurrent fixture
67.4
70.0$0.1930.3k
23Grok 4.5highcurrent fixture
66.9
65.7$0.9748.8k
24Deepseek 4 Prodefaultprevious fixture
64.9
64.7$1.20284.2k
25Terraxhiprevious fixture
63.9
66.9$2.9891.4k
26Solhighcurrent fixture
63.5
62.6$5.8656.1k
27Lunalowcurrent fixture
63.2
64.2$0.1823.7k
28Sollowcurrent fixture
61.9
63.7$5.0227.8k
29Solmedcurrent fixture
61.5
62.5$5.3342.9k
30Lunaxhiprevious fixture
60.2
61.3$0.38140.4k
31Gem 3.5FLdefaultprevious fixture
60.0
59.5$0.55166.2k
32Qwen 3.7Fdefaultprevious fixture
46.4
47.7$0.10337.9k

Scores are percentages, 0–100, graded by code against seeded database state. ▫ marks a row whose eight non-restraint benches ran on the previous fixture: rank it, but do not read a single cell against a ▪ row. Costs come from three different bases — see the run economics below. Output tokens are observed completion tokens for that row's nine-bench run: a workload and verbosity signal, not a quality score.

Not every bench is worth the same

These weights reflect both product importance and how much each bench varies across models. Scaffold carries 3×; defect finding, live edits, schema, and stability carry 2×; restraint and grounded QA carry 1.5×; transcription and reference resolution carry 1×. Grounded QA and reference resolution stay lower because their scores barely move; weights run from 1× to 3× and sum to 16, with observed ranges and the flat mean shown beside the weighted score.

Scaffold roundtrip

build a full spec from a brain-dump

×3

53.199.3

the widest and hardest task and the product's flagship job — and the only bench that genuinely ranks anything, with 27 distinct values across 32 configurations; construction coverage and output restraint are published separately

Proposal gauntlet

exact spec edits via chat proposals

×2

12.587.5

the highest-stakes surface — a wrong proposal mutates a user's live spec — and it still spreads models across 12.5–87.5

Inconsistency needle

find planted spec defects

×2

0.080.0

genuine analytical work and the job people actually buy an analyzer for, but three of its six planted defects are found by everything and three by nothing — 15 of 32 land on the same 66.7, so it cannot carry top weight

False-positive discipline

stay silent on clean specs — re-measured on fixed fixture v3

×1.5

6.7100.0

the paired half of the needle: a defect-finder without restraint is unusable in production, and it spans the widest range on the board (6.7–100)

Schema stress

structured output across surfaces

×2

15.084.8

structured output is the gate every other capability passes through; a model that cannot hold the schema cannot deliver the reasoning behind it

Enumerated data model

transcribe an explicit field list

×1

0.082.0

transcription, not judgement — and ten configurations land on exactly 81.5%

Trap references

resolve near-duplicate entity names

×1

25.0100.0

high-stakes in principle, but 28 of 32 scored exactly 75.0, so it does not get extra weight

Grounded QA

literal answers from the spec, no invention

×1.5

75.0100.0

what users trust the product for daily, but it produces only three distinct values across the whole campaign — 26 of 32 scored 87.5 — so it is close to a constant and is weighted like one

Stability

same analysis five times, scored on consistency

×2

0.080.0

it grades variance rather than raw capability, but this campaign showed variance is the thing that actually bites — 23 distinct values across 32 configurations, and the only bench that repeats its own measurement

The observed range is the lowest and highest score any configuration reached on that bench. A tenth bench, context scaling, was excluded for every model: it never produced a graded case, so it is not on this page and not in any score.

What drives the ranking

The matrix shows every configuration across all 9 benches, starting in leaderboard order. The second header row shows each bench's weight; the final columns show weighted and flat scores. Scaffold and false-positive discipline drive most ranking movement. Trap references and grounded QA stay nearly flat, but they remain because the behaviours matter even when they add little ranking information.

×3×2×2×1.5×2×1×1×1.5×2
Kimi K3defaultcurrent fixture99.375.080.060.084.881.575.087.578.082.080.1
Fable 5lowcurrent fixture93.587.560.086.765.082.0100.087.577.081.482.1
Opus 5xhicurrent fixture98.575.066.773.368.281.575.0100.080.080.779.8
Opus 5medcurrent fixture99.075.066.760.063.681.575.0100.079.078.977.8
Fable 5xhicurrent fixture97.987.560.066.765.081.575.0100.068.078.878.0
Fable 5medcurrent fixture97.775.066.773.363.681.575.087.578.078.677.6
Fable 5highprevious fixture91.787.560.086.762.581.575.087.571.678.578.2
Opus 5lowcurrent fixture99.350.066.766.766.782.075.087.580.075.874.9
Sonnet 5lowcurrent fixture93.950.066.760.066.782.075.0100.080.075.374.9
Sonnet 5medcurrent fixture93.462.566.740.066.781.575.087.580.073.772.6
DeepSeek V4 Flash 0731lowcurrent fixture99.050.066.733.365.082.075.087.580.072.470.9
Terrahighcurrent fixture84.150.060.080.062.574.175.087.578.072.172.4
Terramedcurrent fixture87.350.060.080.065.076.875.087.563.771.471.7
Sonnet 5xhicurrent fixture98.750.066.726.765.081.575.087.577.271.469.8
Opus 5highprevious fixture73.075.066.746.760.081.575.0100.066.770.871.6
DeepSeek V4 Flash 0731maxcurrent fixture99.050.066.760.041.781.975.075.078.370.669.7
Terralowcurrent fixture83.050.060.060.065.074.675.087.575.070.070.0
Sonnet 5highprevious fixture87.062.554.546.762.581.575.087.569.869.869.7
DeepSeek V4 Flash 0731highcurrent fixture96.237.566.720.065.082.075.087.579.069.067.7
Solxhicurrent fixture84.962.566.76.763.672.875.087.578.067.866.4
Lunahighcurrent fixture67.950.066.753.366.772.275.087.578.067.868.6
Lunamedcurrent fixture53.162.560.080.066.776.875.087.568.667.470.0
Grok 4.5highcurrent fixture91.750.060.06.763.682.075.087.574.966.965.7
Deepseek 4 Prodefaultprevious fixture98.537.540.066.738.681.366.787.565.564.964.7
Terraxhiprevious fixture56.750.040.086.758.376.875.087.571.163.966.9
Solhighcurrent fixture86.050.066.76.763.675.575.087.552.763.562.6
Lunalowcurrent fixture62.750.066.780.065.074.250.087.542.163.264.2
Sollowcurrent fixture55.350.054.540.063.672.875.087.574.361.963.7
Solmedcurrent fixture61.050.060.020.063.672.875.087.572.461.562.5
Lunaxhiprevious fixture60.350.054.520.063.671.675.087.569.060.261.3
Gem 3.5FLdefaultprevious fixture85.412.560.093.315.081.025.087.576.060.059.5
Qwen 3.7Fdefaultprevious fixture53.962.50.0100.050.00.075.087.50.046.447.7

Cell shading is the score itself, one hue, no thresholds. Bold is the column best. Every value is a single measurement: the same bench has swung more than thirty points between identical runs elsewhere in this campaign, so treat close cells as a tie. ▫ rows had their eight non-restraint benches measured on the previous fixture — the restraint column was re-measured on the current fixture for everyone, so that one column is comparable across the whole table.

More thinking costs time, not much score

7 models ran 3–4 reasoning efforts each. Going from a model's cheapest setting to its dearest improved the weighted score by at most 6.0 points, and 5 of 7 got worse, while runtime grew 2.04.1×. For 5 of 7, the most expensive setting is not the best one. A full rerun flipped one model's effort order too, so the effect is small and unstable; high and xhigh have not shown a repeatable benefit over medium here. Other workloads are untested.

DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731: weighted score and wall clock at each reasoning effort
low current fixture72.421.0m
high current fixture69.040.5m
max current fixture70.641.3m

−1.8 quality · 2.0× time

Fable 5

Fable 5: weighted score and wall clock at each reasoning effort
low current fixture81.411.9m
med current fixture78.617.6m
high previous fixture78.522.9m
xhi current fixture78.831.2m

−2.6 quality · 2.6× time

Opus 5

Opus 5: weighted score and wall clock at each reasoning effort
low current fixture75.813.4m
med current fixture78.920.0m
high previous fixture70.830.4m
xhi current fixture80.735.2m

+4.9 quality · 2.6× time

Sonnet 5

Sonnet 5: weighted score and wall clock at each reasoning effort
low current fixture75.313.1m
med current fixture73.725.1m
high previous fixture69.830.6m
xhi current fixture71.448.5m

−4.0 quality · 3.7× time

Luna

Luna: weighted score and wall clock at each reasoning effort
low current fixture63.211.3m
med current fixture67.414.9m
high current fixture67.830.2m
xhi previous fixture60.246.5m

−3.0 quality · 4.1× time

Terra

Terra: weighted score and wall clock at each reasoning effort
low current fixture70.011.7m
med current fixture71.413.9m
high current fixture72.114.8m
xhi previous fixture63.932.3m

−6.1 quality · 2.8× time

Sol

Sol: weighted score and wall clock at each reasoning effort
low current fixture61.914.3m
med current fixture61.534.0m
high current fixture63.524.9m
xhi current fixture67.834.8m

+6.0 quality · 2.4× time

Thick bar: weighted score on a 0–100 scale. Thin bar underneath: wall clock, scaled against the longest run on the board (48.5 minutes). The footer compares the lowest effort with the highest. Read only ▪-to-▪ pairs as a clean effort effect — a ▫ row had its eight non-restraint benches measured on the previous fixture. Wall clock is operational time, not model latency: the hosted model group is throttled and Anthropic/OpenAI runs pay process-spawn overhead per call.

The best team mixes strengths

An ideal router searches every team built from 31 eligible configurations and gives each bench to the team's strongest member. Qwen 3.7 Flash is excluded because its single 100% restraint result would skew the search; the next-best restraint configuration fills that slot. The best 3-model team reaches 89%, up from a solo baseline of 82%. That is a 7.1-point gain. The pale line marks solo performance; cyan shows what routing adds.

  • 1 model

    82%

    • Kimi K3defaultcurrent fixture
  • 2 models

    87.6%+5.66

    • Kimi K3defaultcurrent fixture
    • Fable 5lowcurrent fixture
  • 3 models

    89%+7.08

    • Kimi K3defaultcurrent fixture
    • Fable 5lowcurrent fixture
    • Opus 5xhicurrent fixture

Three models, one team

The best three-model team combines:

  • Kimi K3defaultcurrent fixture
  • Fable 5lowcurrent fixture
  • Opus 5xhicurrent fixture

Fable 5 · low earns its place as a complement to the other two, not as a standalone recommendation. This is a perfect-router ceiling: it assumes every task reaches the strongest member, and a different workload could produce a different trio.

For the best 3-model team, the member that wins each bench
Scaffold roundtripKimi K3defaultcurrent fixture99.3×3
Proposal gauntletFable 5lowcurrent fixture87.5×2
Inconsistency needleKimi K3defaultcurrent fixture80.0×2
False-positive disciplineFable 5lowcurrent fixture86.7×1.5
Schema stressKimi K3defaultcurrent fixture84.8×2
Enumerated data modelFable 5lowcurrent fixture82.0×1
Trap referencesFable 5lowcurrent fixture100.0×1
Grounded QAOpus 5xhicurrent fixture100.0×1.5
StabilityOpus 5xhicurrent fixture80.0×2

This is a ceiling, not a result. It assumes a router that always picks correctly — nothing here demonstrates such a router, and building one is its own hard problem. Every cell behind it is a single measurement with real run-to-run variance. Treat 89% as the headroom a routing strategy could chase, not a score any deployment has achieved.

185× the cost buys 6.4 more points

Each dot is one full suite. Cost runs horizontally on a log scale; weighted score runs vertically. Hover any dot to see which reasoning effort it was. The connected lines walk one model up its own effort ladder, and their flatness is how little extra score spending more buys. DeepSeek V4 Flash 0731 · low costs $0.09 for 72.4%; Fable 5 · xhigh costs $16.61 for 78.8%. The cheapest row is DeepSeek V4 Flash 0731 · low at $0.09, ranked 11th of 32.

  • The rest
  • Anthropic
  • OpenAI
405060708090$0.10$0.30$1$3$10Cost of one nine-bench suite (log scale)Weighted scoreDeepSeek V4 Flash 0731Qwen 3.7FLunaGem 3.5FLGrok 4.5Deepseek 4 ProTerraKimi K3Sonnet 5SolOpus 5Fable 5

Dot size is reasoning effort — smallest low, largest xhigh; the hosted provider models ran at their provider's default and appear once. Costs are shown as the dollar total for each suite. OpenAI's models actually bill to a subscription, so their dollars answer what the same work would have cost at published usage rates, not what was charged.

What the whole campaign cost

One line per suite: every figure is summed from the per-call telemetry in that run's report, so a model measured at four settings appears four times, once for each. Read these as what a single nine-bench suite cost, not as a price list — the spread runs from about ten cents to seventeen dollars, and almost all of it is what each model charges per token rather than anything the benchmark did. Prompt tokens are the one figure the reports only aggregate per model, so this table shows completions.

  • The rest
  • Anthropic
  • OpenAI
Qwen 3.7 Flashdefault0.34M$0.1024.4m
DeepSeek V4 Flash 0731low0.17M$0.0921.0m
DeepSeek V4 Flash 0731high0.30M$0.1340.5m
DeepSeek V4 Flash 0731max0.32M$0.1441.3m
DeepSeek V4 Prodefault0.28M$1.2035.2m
Gemini 3.5 Flash-Litedefault0.17M$0.558.1m
Grok 4.5high0.05M$0.9724.7m
Moonshot Kimi K3default0.10M$2.2797.6m
Claude Fable 5low0.06M$11.3211.9m
Claude Fable 5med0.09M$12.6217.6m
Claude Fable 5high0.12M$13.6722.9m
Claude Fable 5xhi0.17M$16.6131.2m
Claude Opus 5low0.07M$6.0113.4m
Claude Opus 5med0.10M$6.9420.0m
Claude Opus 5high0.16M$7.9930.4m
Claude Opus 5xhi0.19M$8.6435.2m
Claude Sonnet 5low0.08M$4.0613.1m
Claude Sonnet 5med0.17M$5.7425.1m
Claude Sonnet 5high0.19M$5.6330.6m
Claude Sonnet 5xhi0.24M$8.0848.5m
GPT-5.6 Lunalow0.02M$0.1811.3m
GPT-5.6 Lunamed0.03M$0.1914.9m
GPT-5.6 Lunahigh0.08M$0.2630.2m
GPT-5.6 Lunaxhi0.14M$0.3846.5m
GPT-5.6 Terralow0.03M$1.9811.7m
GPT-5.6 Terramed0.03M$1.9713.9m
GPT-5.6 Terrahigh0.03M$2.0414.8m
GPT-5.6 Terraxhi0.09M$2.9832.3m
GPT-5.6 Sollow0.03M$5.0214.3m
GPT-5.6 Solmed0.04M$5.3334.0m
GPT-5.6 Solhigh0.06M$5.8624.9m
GPT-5.6 Solxhi0.08M$6.6234.8m
Campaign32 suites4.04M$145.5714.6h

Wall clock is scored-bench time, an operational number rather than model latency: the hosted model group is throttled to 15 requests a minute, while Anthropic and OpenAI runs are unthrottled but pay process-spawn overhead on every call. One row to read with care — DeepSeek V4 Flash 0731 is the cheapest full suite at $0.09 and scores 72.4%; low cost is not the same as broad capability.

We went looking for a benchmark that measured this. There wasn't one.

Reading and editing are different skills

Public leaderboards measure whether a model can write a function. Ours asks whether it can change a living specification without quietly breaking it. The two skills barely know each other.

Specs are held together by references

A flow step points at a role; a field points at its parent. Rename the wrong near-identical entity and nothing turns red — the damage is silent, and usually found weeks later by whoever tries to build from it.

Doing nothing looks like doing it right

A model that proposed nothing and a model that finished the job read exactly the same in a chat window. So we stopped taking anyone's word for it and started checking the database afterwards.

Nothing on this page was graded by another model. Each bench seeds its own workspace in an isolated database, runs against a fingerprinted fixture, and is scored by plain code — set overlap, precision and recall, or simply reading the project afterwards to see whether the change is there. A grader that cannot be charmed is a harsh one, which is why the scores are lower than you might expect.

Before you quote us on any of this

A benchmark tells you about the shape of a problem long before it tells you anything reliable about a particular model. We would rather point at the thin spots ourselves than have you find them, so here they are, in the order that would embarrass us most.

More models are in the queue — this page updates when their runs land.

We ran Sol's whole suite twice

2 passes · 3 efforts · same fixture

Every other model here was measured once. Sol we measured twice — three complete suites, then three more a day later, same benches, same lane, nothing changed but the clock. Two of the three came back within about a point of where they started. The third moved by nearly seven, which was enough to change which setting looked best.

That is the closest thing to an error bar this page has, and it is why we would rather you read the ranking as a shape than a verdict. Nothing was wrong with either pass — this is simply what one measurement is worth.

Sol's first and second measurement of every bench, at each of three reasoning efforts, with the change between them.
Benchmediumhighxhigh
1st → 2ndΔ1st → 2ndΔ1st → 2ndΔ
Scaffold roundtrip55.361.0+5.784.486.0+1.655.484.9+29.5
Proposal gauntlet50.050.062.550.0−12.550.062.5+12.5
Inconsistency needle54.560.0+5.560.066.7+6.761.566.7+5.2
False-positive discipline40.020.0−20.06.76.740.06.7−33.3
Schema stress62.563.6+1.163.663.637.563.6+26.1
Enumerated data model72.372.8+0.572.375.5+3.272.372.8+0.5
Trap references75.075.075.075.075.075.0
Grounded QA87.587.587.587.587.587.5
Stability74.972.4−2.576.452.7−23.769.478.0+8.6

Both passes are real provider runs through the product, graded by the same deterministic code (coverage-v1 + restraint-v1 for the scaffold column). The leaderboard shows the second pass. We publish the first rather than quietly replacing it, because the gap between them is the most useful number on this page.

Sol kept finding problems in clean controls

6.7 · 16 findings

This is the archived false-positive-discipline run for Sol at high effort. The bench treats all three controls as clean, but Sol returned findings on every one. The score is the average of the three control scores: 0, 0, 20 6.7.

Base spec

0.0 · 6

Findings returned against a control the bench marked clean.

Near-duplicate controls

0.0 · 6

Findings returned against a control the bench marked clean.

Hostile-content controls

20.0 · 4

Findings returned against a control the bench marked clean.

This is a compact summary of the archived output, not a new model run. Sol mostly raised workflow consistency and handoff concerns; the full finding text stays in the report rather than on this page. Source: bench-parallel-2026-07-30T20-23-32-828Z.json.

  • One pass per bench — and we measured what that costs

    Every score here is a single measurement, with no confidence interval behind it. So we tested what that is worth: we re-ran one model's three efforts end to end, all twenty-seven cells. Nine came back identical to the decimal. Four moved by more than twenty points. And the ranking of the three efforts inverted completely — the setting that had finished last finished first. The weighted totals held up far better than the cells inside them, which is the strongest argument on this page for reading the weighted column and ignoring any single bench. Treat a gap of a point or two as a tie.

  • The weights are ours

    Weights run 1× to 3× and are our editorial judgement about which benches represent real product work, not an output of the harness. They were revised once: grounded QA and inconsistency needle originally carried 3× each, until we counted how little they separate anything — three and six distinct values across 32 configurations — and moved that weight to scaffold and stability, which do. Re-weighting turned out to be nearly cosmetic, because a bench everyone ties on shifts everyone together; it changed no leader and moved nobody more than four places. We publish the flat unweighted mean next to the weighted one so you can see exactly what our judgement did to the ordering — and disagree with it.

  • Cost is a rough comparison

    The cost column shows a full-suite dollar total. Treat it as a rough comparison rather than an invoice, especially for OpenAI's subscription-billed models.

  • CrystalBench is ours, for now

    It runs against CrystalSpec's own AI surfaces and private fixtures, so it isn't something you can point at your own stack today. The methodology, though — that part you're very welcome to steal.

Come build a spec worth benchmarking

Every fixture behind CrystalBench is just an ordinary CrystalSpec project — typed flows, roles, data models, and test cases, all versioned and reviewable. Build one of your own and see what your coding agents make of it.

Start free