The Everyday Model IndexThe cheapest model you can trust by default.
Which model configuration should your AI agents run by default? We call that configuration the everyday model, and this index picks it with public benchmarks and one costing method.
The decision
What should be your default model?
The everyday model is the configuration that handles the nine in ten prompts that finish quickly. The top tier is held back for tasks that are genuinely hard or long.
Why it matters
9 in 10 tasks do not need the top-tier model.
94% of coding prompts and 92% of knowledge-work prompts finish inside 15 minutes and carry 71% and 50% of token spend. On those prompts a correct answer costs $0.14 on the cheapest qualifier against $2.04 on Claude Fable 5.1 at max, a 15× difference at 86% efficacy against 98%. That is the saving the default decides.
How we judged
Honest first, capable second, then cheapest.
A confident wrong answer costs labor hours, so honesty is a gate. Then at least 85% efficacy on a broad set of PhD-level, rigorous evaluations. Then rank by cost per correct answer.
Most work is short
Across our own agent telemetry, more than nine in ten prompts resolve inside fifteen minutes. These are the mundane and moderately complex tasks that make up a working day. Run them on the everyday model, the cheapest configuration you can trust, and hold the top tier for the rest.
94%
of coding-agent prompts resolve in under 15 minutes
17,609 measured prompts
92%
of knowledge-work prompts resolve in under 15 minutes
3,796 measured prompts
Where the saving is, and where it is not
Those sub-15-minute prompts carry 71% of coding spend and 50% of knowledge-work spend, and that is where the everyday model pays: a correct answer costs $0.14 on the cheapest qualifier against $2.04 on Claude Fable 5.1 at max, a 15× difference at 86% efficacy against 98%. Start every prompt there and the majority of the bill moves to the cheap end.
The caveat is the tail. It is small by count and large by cost, larger in knowledge work than in coding, so an everyday model is not a licence to stop paying attention to it. The escalation rule below handles it. Route on predicted duration, not on prompt count, or the saving will disappoint.
Zaun internal data · August 2026 · a prompt is measured from submission to its last completion event
Three questions
Every configuration is asked the same three, in this order. Fail the first and cost never matters; fail the second and it does not reach the ranking.
First
Is it honest?
Right answers minus confidently wrong ones. Below zero it asserts more than it knows. Confident errors are paid for in labor hours, not tokens, so a model that asserts more than it knows is out before cost is even considered.
Second
Is it capable enough?
How close it gets to the best score on each test, averaged across six benchmarks. The six are graduate and research-level science, frontier exams, scientific coding, terminal work and long-document reasoning. A configuration that holds 85% of the best score across them is more than capable enough for everyday work; that is the bar to enter the ranking.
Third
What does it cost?
What one correct answer costs once you have paid for the attempts that failed. This is what the ranking sorts on.
The ranking
Clear the honesty gate by more than measurement noise, keep at least 85% efficacy, then rank by cost per solve. Cheapest first. A configuration is a model plus a reasoning-effort setting. GPT-5.6 Sol at medium and GPT-5.6 Sol at max are two configurations of one model, and they rank differently. Open any row for its per-benchmark detail and its weak spot. Every configuration measured is here, including the ones that do not qualify and why.
#ConfigurationCost per solveCost per solveWhat one correct answer costs once you have paid for the attempts that failed.EfficacyEfficacyHow close it gets to the best score on each test, averaged across six benchmarks.HonestyHonestyRight answers minus confidently wrong ones. Below zero it asserts more than it knows.
1GPT-5.6 Solmedium$0.1486%passes›
Weak spotCritPt at 71%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+19.4per 100 questions, right answers minus confident errors
Cost per fact$0.05What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
92.6%
96%
$0.0157
$0.017
HLE
42.2%
71%
$0.0430
$0.102
CritPt
22.9%
71%
$0.1259
$0.550
SciCode
56.5%
91%
$0.0285
$0.050
Terminal-Bench
86.1%
94%
$0.2302
$0.267
Long context
74.3%
89%
$0.3830
$0.515
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
2GPT-6 Astralow$0.1788%passes›
Weak spotCritPt at 81%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+40.5per 100 questions, right answers minus confident errors
Cost per fact$0.02What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.1%
97%
$0.0155
$0.017
HLE
49.2%
83%
$0.0389
$0.079
CritPt
26.3%
81%
$0.1449
$0.551
SciCode
51.0%
82%
$0.0425
$0.083
Terminal-Bench
88.0%
96%
$0.2903
$0.330
Long context
72.7%
87%
$0.9501
$1.308
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
3GPT-5.6 Solhigh$0.1989%passes›
Weak spotHLE at 78%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+20.4per 100 questions, right answers minus confident errors
Cost per fact$0.09What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
92.8%
96%
$0.0243
$0.026
HLE
46.0%
78%
$0.0828
$0.180
CritPt
25.7%
80%
$0.2453
$0.955
SciCode
56.9%
92%
$0.0353
$0.062
Terminal-Bench
87.3%
96%
$0.2794
$0.320
Long context
75.3%
90%
$0.3859
$0.512
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
4Muse Spark 1.3xhigh$0.2190%passes›
Weak spotHLE at 80%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+23.1per 100 questions, right answers minus confident errors
Cost per fact$0.07What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
94.1%
98%
$0.0298
$0.032
HLE
47.5%
80%
$0.1041
$0.219
CritPt
26.0%
81%
$0.3080
$1.185
SciCode
58.6%
95%
$0.0371
$0.063
Terminal-Bench
85.4%
93%
$0.9380
$1.099
Long context
79.3%
95%
$0.1315
$0.166
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
5Gemini 3.8 Flashhigh$0.2286%passes›
Weak spotCritPt at 57%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+29.6per 100 questions, right answers minus confident errors
Cost per fact$0.04What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
95.3%
99%
$0.0245
$0.026
HLE
47.8%
81%
$0.0920
$0.192
CritPt
18.3%
57%
$0.2800
$1.531
SciCode
53.6%
86%
$0.0769
$0.143
Terminal-Bench
87.6%
96%
$0.7178
$0.819
Long context
81.0%
97%
$0.0988
$0.122
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
6GPT-6 Astramedium$0.2691%passes›
Weak spotSciCode at 82%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+42.2per 100 questions, right answers minus confident errors
Cost per fact$0.04What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.9%
98%
$0.0244
$0.026
HLE
52.7%
89%
$0.1005
$0.191
CritPt
29.1%
90%
$0.2851
$0.978
SciCode
51.0%
82%
$0.0509
$0.100
Terminal-Bench
89.5%
98%
$0.4506
$0.503
Long context
76.3%
92%
$0.9533
$1.249
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
7GPT-5.6 Solxhigh$0.2691%passes›
Weak spotHLE at 80%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+21.0per 100 questions, right answers minus confident errors
Cost per fact$0.16What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.1%
97%
$0.0370
$0.040
HLE
47.3%
80%
$0.1421
$0.300
CritPt
28.6%
89%
$0.4367
$1.527
SciCode
56.0%
90%
$0.0449
$0.080
Terminal-Bench
89.5%
98%
$0.3821
$0.427
Long context
76.3%
92%
$0.3876
$0.508
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
8GPT-6 Astrahigh$0.3692%passes›
Weak spotSciCode at 83%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+43.7per 100 questions, right answers minus confident errors
Cost per fact$0.06What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
94.9%
99%
$0.0371
$0.039
HLE
53.1%
90%
$0.1618
$0.305
CritPt
28.9%
89%
$0.4335
$1.502
SciCode
51.6%
83%
$0.0696
$0.135
Terminal-Bench
89.9%
98%
$0.6450
$0.718
Long context
76.0%
91%
$0.9597
$1.263
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
9Claude Opus 5medium$0.3889%passes›
Weak spotSciCode at 82%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+31.0per 100 questions, right answers minus confident errors
Cost per fact$0.02What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
91.9%
96%
$0.0241
$0.026
HLE
51.3%
87%
$0.1870
$0.364
CritPt
26.9%
83%
$1.4033
$5.217
SciCode
50.7%
82%
$0.0369
$0.073
Terminal-Bench
86.1%
94%
$0.7522
$0.874
Long context
78.7%
94%
$0.7378
$0.938
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
10Claude Fable 5.1low$0.4090%passes›
Weak spotHLE at 83%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+34.1per 100 questions, right answers minus confident errors
Cost per fact$0.03What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
88.1%
92%
$0.0164
$0.019
HLE
48.9%
83%
$0.1238
$0.253
CritPt
27.7%
86%
$1.2559
$4.534
SciCode
55.7%
90%
$0.0690
$0.124
Terminal-Bench
85.0%
93%
$0.6832
$0.804
Long context
79.3%
95%
$1.4690
$1.853
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
11 more configurations qualify11GPT-5.6 Solmax$0.4194%passes›
Weak spotHLE at 84%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+22.0per 100 questions, right answers minus confident errors
Cost per fact$0.39What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
94.1%
98%
$0.0732
$0.078
HLE
49.5%
84%
$0.2684
$0.542
CritPt
32.3%
100%
$0.8591
$2.660
SciCode
56.1%
91%
$0.0797
$0.142
Terminal-Bench
88.0%
96%
$0.5435
$0.618
Long context
77.7%
93%
$0.4016
$0.517
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
12Kimi K3maxopen$0.5089%passes›
Weak spotCritPt at 72%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+19.7per 100 questions, right answers minus confident errors
Cost per fact$0.67What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.5%
97%
$0.1512
$0.162
HLE
46.9%
79%
$0.3936
$0.839
CritPt
23.4%
72%
$0.9019
$3.854
SciCode
58.7%
95%
$0.0602
$0.103
Terminal-Bench
85.0%
93%
$0.6271
$0.738
Long context
82.7%
99%
$0.3086
$0.373
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
13Claude Fable 5.1medium$0.5192%passes›
Weak spotSciCode at 89%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+37.6per 100 questions, right answers minus confident errors
Cost per fact$0.04What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
88.6%
92%
$0.0263
$0.030
HLE
53.8%
91%
$0.2254
$0.419
CritPt
29.1%
90%
$1.8039
$6.199
SciCode
55.3%
89%
$0.0756
$0.137
Terminal-Bench
88.0%
96%
$0.7691
$0.874
Long context
78.7%
94%
$1.4752
$1.875
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
14Claude Opus 5high$0.5592%passes›
Weak spotSciCode at 88%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+33.7per 100 questions, right answers minus confident errors
Cost per fact$0.03What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.7%
97%
$0.0452
$0.048
HLE
52.8%
89%
$0.3521
$0.667
CritPt
28.3%
88%
$1.9945
$7.048
SciCode
54.3%
88%
$0.0481
$0.089
Terminal-Bench
87.6%
96%
$1.2258
$1.399
Long context
76.3%
92%
$0.7456
$0.977
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
15GPT-6 Astraxhigh$0.5593%passes›
Weak spotSciCode at 80%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+43.4per 100 questions, right answers minus confident errors
Cost per fact$0.12What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
96.3%
100%
$0.0726
$0.075
HLE
54.6%
92%
$0.2532
$0.464
CritPt
31.4%
97%
$0.7980
$2.539
SciCode
49.5%
80%
$0.1216
$0.245
Terminal-Bench
89.1%
98%
$0.8704
$0.977
Long context
74.0%
89%
$0.9699
$1.311
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
16Claude Fable 5.1high$0.7394%passes›
Weak spotLong context at 92%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+40.8per 100 questions, right answers minus confident errors
Cost per fact$0.05What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
90.6%
94%
$0.0553
$0.061
HLE
55.9%
95%
$0.3970
$0.710
CritPt
30.3%
94%
$2.7477
$9.068
SciCode
57.6%
93%
$0.0891
$0.155
Terminal-Bench
89.9%
98%
$1.1876
$1.321
Long context
77.0%
92%
$1.4810
$1.923
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
17Claude Opus 5xhigh$0.7392%passes›
Weak spotCritPt at 86%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+35.4per 100 questions, right answers minus confident errors
Cost per fact$0.04What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.7%
97%
$0.0635
$0.068
HLE
54.4%
92%
$0.5171
$0.951
CritPt
27.7%
86%
$2.4581
$8.874
SciCode
55.0%
89%
$0.0723
$0.132
Terminal-Bench
88.0%
96%
$1.8677
$2.122
Long context
76.3%
92%
$0.7529
$0.987
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
18GPT-6 Astramax$0.7694%passes›
Weak spotSciCode at 87%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+43.4per 100 questions, right answers minus confident errors
Cost per fact$0.22What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
96.1%
100%
$0.0924
$0.096
HLE
54.7%
93%
$0.3781
$0.691
CritPt
31.7%
98%
$1.1662
$3.677
SciCode
54.1%
87%
$0.2323
$0.430
Terminal-Bench
88.4%
97%
$1.1997
$1.357
Long context
74.3%
89%
$0.9973
$1.342
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
19Claude Opus 5max$0.9493%passes›
Weak spotSciCode at 90%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+37.1per 100 questions, right answers minus confident errors
Cost per fact$0.07What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.2%
97%
$0.0921
$0.099
HLE
54.9%
93%
$0.6736
$1.227
CritPt
29.1%
90%
$2.8845
$9.912
SciCode
55.7%
90%
$0.1177
$0.211
Terminal-Bench
89.1%
98%
$2.4213
$2.717
Long context
75.7%
91%
$0.7611
$1.005
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
20Claude Fable 5.1xhigh$1.3297%passes›
Weak spotLong context at 94%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+42.4per 100 questions, right answers minus confident errors
Cost per fact$0.18What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.4%
97%
$0.0939
$0.101
HLE
58.7%
99%
$1.0254
$1.747
CritPt
31.1%
96%
$4.5338
$14.578
SciCode
60.1%
97%
$0.2410
$0.401
Terminal-Bench
91.0%
100%
$2.4726
$2.717
Long context
78.0%
94%
$1.5078
$1.933
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
21Claude Fable 5.1max$2.0498%passes›
Weak spotCritPt at 92%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+43.5per 100 questions, right answers minus confident errors
Cost per fact$0.62What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.7%
97%
$0.1378
$0.147
HLE
59.1%
100%
$1.5859
$2.683
CritPt
29.7%
92%
$5.7138
$19.238
SciCode
62.0%
100%
$0.6585
$1.062
Terminal-Bench
91.4%
100%
$4.1512
$4.542
Long context
80.0%
96%
$1.5845
$1.981
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
20 are honest but fall under the 85% efficacy bar
20 of these 20 are held back by the same benchmark: CritPt, research-level physics. Most clear the other five comfortably. Gemini 3.7 Flash at high (84%), Muse Spark 1.2 at xhigh (84%), Grok 4.6 at xhigh (84%) sit within two points of the bar, so a single benchmark is deciding whether they qualify.
Weak spotCritPt at 62%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+4.3per 100 questions, right answers minus confident errors
Cost per fact$0.65What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
93.5%
97%
$0.0658
$0.070
HLE
42.4%
72%
$0.2016
$0.475
CritPt
20.0%
62%
$0.4828
$2.414
SciCode
51.6%
83%
$0.1006
$0.195
Terminal-Bench
82.0%
90%
$0.6941
$0.846
Long context
75.3%
90%
$0.2220
$0.295
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
·Qwen3.8 Maxdefault82% efficacy$0.4582%passes›
Weak spotCritPt at 62%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+3.4per 100 questions, right answers minus confident errors
Cost per fact$0.84What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
92.7%
96%
$0.0735
$0.079
HLE
43.0%
73%
$0.2234
$0.519
CritPt
20.0%
62%
$0.5092
$2.546
SciCode
52.9%
85%
$0.1106
$0.209
Terminal-Bench
81.3%
89%
$1.0157
$1.250
Long context
74.3%
89%
$0.2218
$0.298
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
·Claude Sonnet 5max81% efficacy$1.3381%passes›
Weak spotCritPt at 52%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+16.4per 100 questions, right answers minus confident errors
Cost per fact$0.18What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
91.1%
95%
$0.3663
$0.402
HLE
41.3%
70%
$1.0208
$2.472
CritPt
16.9%
52%
$2.3138
$13.691
SciCode
53.6%
87%
$0.1917
$0.358
Terminal-Bench
80.5%
88%
$1.8879
$2.345
Long context
77.0%
92%
$0.3679
$0.478
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
6 are too close to call on honesty
Their net knowledge score sits within measurement noise of zero, so the test cannot say whether they are right more often than confidently wrong. They are not ranked, because a cost per correct answer needs a denominator we can stand behind. The closest case is GPT-5.6 Terra at max, which nets +0.05 against a standard error of about 1.29, so it is 0.04 standard errors from zero.
Weak spotCritPt at 10%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−0.4per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
86.7%
90%
$0.0396
$0.046
HLE
28.4%
48%
$0.0996
$0.351
CritPt
3.1%
10%
$0.2226
$7.082
SciCode
39.9%
64%
$0.0124
$0.031
Terminal-Bench
53.9%
59%
$1.1414
$2.116
Long context
71.0%
85%
$0.0845
$0.119
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
·GPT-5.6 Terramaxhonesty +0.05$0.3890%unclear›
Weak spotHLE at 73%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+0.1per 100 questions, right answers minus confident errors
Cost per fact$140.60What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
92.5%
96%
$0.0730
$0.079
HLE
42.9%
73%
$0.2277
$0.531
CritPt
30.0%
93%
$0.6084
$2.028
SciCode
53.9%
87%
$0.1144
$0.212
Terminal-Bench
88.0%
96%
$0.5732
$0.651
Long context
79.7%
96%
$0.2130
$0.267
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
20 fail the honesty gate at every effort setting
These are right less often than they are confidently wrong. Several are also the cheapest configurations measured: GPT-5.6 Luna (low) would top the ranking at $0.013 per solve, roughly 10x below the cheapest qualifier. They are excluded anyway, because every confident error is undone by a person at a labor rate that swamps the saving. Note that they look fine on a standard test: their median GPQA Diamond score is 90%.
·GPT-5.6 Lunalowhonesty −15$0.01355%fails›
Weak spotCritPt at 8%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−14.7per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
83.5%
87%
$0.0008
$0.001
HLE
19.8%
34%
$0.0014
$0.007
CritPt
2.6%
8%
$0.0037
$0.142
SciCode
45.6%
74%
$0.0015
$0.003
Terminal-Bench
43.4%
48%
$0.0239
$0.055
Long context
65.3%
78%
$0.0191
$0.029
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
·GPT-5.6 Lunamediumhonesty −13$0.01561%fails›
Weak spotCritPt at 15%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−13.2per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
85.9%
89%
$0.0013
$0.002
HLE
25.8%
44%
$0.0029
$0.011
CritPt
4.9%
15%
$0.0066
$0.135
SciCode
45.8%
74%
$0.0018
$0.004
Terminal-Bench
53.2%
58%
$0.0243
$0.046
Long context
72.0%
86%
$0.0193
$0.027
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
Weak spotCritPt at 31%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−10.2per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
89.6%
93%
$0.0396
$0.044
HLE
35.0%
59%
$0.1047
$0.299
CritPt
10.0%
31%
$0.2563
$2.563
SciCode
47.5%
77%
$0.0238
$0.050
Terminal-Bench
67.4%
74%
$0.4143
$0.615
Long context
75.0%
90%
$0.0969
$0.129
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
·Qwen3.8 27Bxhighopenhonesty −10$0.3270%fails›
Weak spotCritPt at 17%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−10.0per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
90.5%
94%
$0.0464
$0.051
HLE
33.9%
57%
$0.1326
$0.391
CritPt
5.4%
17%
$0.3217
$5.926
SciCode
44.7%
72%
$0.0645
$0.144
Terminal-Bench
79.8%
87%
$0.5963
$0.748
Long context
77.3%
93%
$0.0617
$0.080
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
4 never solve one of the benchmarks at all·Gemini 3.5 Flash-Litedefaultscores zero on CritPtno price56%passes›
Weak spotCritPt at 0%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+5.2per 100 questions, right answers minus confident errors
Cost per fact$0.06What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
83.8%
87%
$0.0122
$0.015
HLE
18.8%
32%
$0.0271
$0.144
CritPt
0.0%
0%
$0.0704
never solves
SciCode
40.9%
66%
$0.0126
$0.031
Terminal-Bench
53.6%
59%
$0.2571
$0.480
Long context
74.7%
90%
$0.0390
$0.052
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
·Claude Haiku 4.5reasoningscores zero on CritPtno price49%fails›
Weak spotCritPt at 0%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−4.4per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
67.2%
70%
$0.0749
$0.112
HLE
10.4%
18%
$0.1245
$1.197
CritPt
0.0%
0%
$0.1648
never solves
SciCode
43.3%
70%
$0.0507
$0.117
Terminal-Bench
44.2%
48%
$0.6990
$1.581
Long context
73.7%
88%
$0.1213
$0.165
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
·Qwen3.8 27Bmediumopenscores zero on CritPtno price56%fails›
Weak spotCritPt at 0%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−36.1per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
84.5%
88%
$0.0152
$0.018
HLE
14.1%
24%
$0.0390
$0.277
CritPt
0.0%
0%
$0.0918
never solves
SciCode
38.1%
61%
$0.0117
$0.031
Terminal-Bench
65.2%
71%
$0.4084
$0.627
Long context
76.3%
92%
$0.0576
$0.075
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
·Qwen3.8 27Blowopenscores zero on CritPtno price57%fails›
Weak spotCritPt at 0%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−26.7per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
Benchmark
Score
Share of best
Per attempt
Per solve
GPQA
84.5%
88%
$0.0141
$0.017
HLE
14.0%
24%
$0.0217
$0.154
CritPt
0.0%
0%
$0.0509
never solves
SciCode
39.8%
64%
$0.0082
$0.021
Terminal-Bench
67.4%
74%
$0.4303
$0.638
Long context
74.7%
90%
$0.0571
$0.076
The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.
Cheapest qualifying configuration $0.14 · most expensive $2.04 · a 15× range for 12 points of efficacy
What the data shows
Four findings, each with the point first and the evidence behind it one click down.
3.3×Use the reasoning lever, but keep the big model.›
Claude Sonnet 5 at max effort costs less per token than Claude Fable 5.1, and 3.3 times more per correct answer, because it uses more tokens and solves fewer problems per attempt. Fable 5.1 at low effort is cheaper and 9 points more capable. The sticker price told the opposite story.
Sonnet 5 max $1.33 per solve at 81% efficacy · Fable 5.1 low $0.40 at 90% · same vendor, both pass the honesty gate
20 of 71The cheapest configurations are among the ones that make things up.›
20 configurations are right less often than they are confidently wrong. GPT-5.6 Luna (low) would top the whole ranking at $0.013 per solve if the gate did not exist. It is excluded because every confident error is undone by a person, at a labor rate that swamps the saving. These models look fine on a standard test: their median GPQA Diamond score is 90%. Honesty is not something you can read off accuracy.
Green passes, amber is within measurement noise of zero, red fails · every GPT-5.6 Luna setting fails, best margin −10.3
3×The effort setting moves cost more than the model name does.›
GPT-5.6 Sol alone spans $0.14 to $0.41 per solve and 86% to 94% efficacy across its effort settings. That is most of the useful range without changing model. The same is true of every frontier family here, which is why a configuration, model plus effort, is the unit ranked and never the model alone.
Sol at low (78%) falls under the bar; at non-reasoning its honesty is too close to call · the setting matters on both gates
1 of 10Labs outside the big three hold top-ten places.›
Muse Spark 1.3 ranks 4. Of the 21 configurations that qualify, one has open weights and can be self-hosted. The field is wider than the three names most buying conversations start with, and it is moving fast: GPT-6 Astra, released 3 September 2026, was added on 4 September 2026 and holds 3 of the top 10 places.
Open-weight rows are marked on the ranking; their price is the creator's own rate, which third-party hosts routinely undercut
When to leave it
Three triggers. Everything else stays on the everyday model.
1The reasoning is genuinely hard.›
CritPt, research-level physics, still separates tiers sharply. Two cheap configurations elsewhere on the board score exactly zero on it while still billing for the attempt. Passing GPQA Diamond does not predict passing CritPt, so a high headline score is not a licence to send it your hardest work.
2The model must state facts it cannot look up.›
That is what the honesty column measures. Plus 31 means that across 100 questions the model finished 31 correct answers ahead of its confident mistakes. Minus 10.3 means it finished about 10 confident mistakes behind. Every GPT-5.6 Luna configuration sits below zero, so none qualifies at any effort. Put retrieval in front of the model and the gate relaxes, because it is no longer asserting from memory.
3The task will run long.›
Fewer than one prompt in ten runs past 15 minutes, but those prompts carry 29% of coding spend and 50% of knowledge-work spend. This index scores bounded work only and does not adjudicate what to escalate to. Route on predicted duration, not on prompt count.
How we measured
Public benchmarks, one costing method throughout, and a rule stated in advance rather than fitted to the answer.
Abstract
We rank 71 model configurations from 30 models and 11 labs on the cost of one correct answer, using seven public component evaluations of the Artificial Analysis Intelligence Index and first-party prices, retrieved 1 September 2026, with GPT-6 Astra from pages retrieved 4 September 2026. A configuration enters the ranking only if it clears a knowledge-honesty gate by more than measurement noise and holds at least 85% of the best score across six accuracy evaluations; 21 do. Cost per solve is cost per attempt divided by share solved, combined by geometric mean so no single benchmark dominates. We publish the sensitivity of the qualifying set to the bar and to the benchmark set, and every figure traces to a dated snapshot and versioned scripts.
1
Scope and inclusion
Every model on the Artificial Analysis leaderboard carrying a measured Intelligence Index score and a score on all seven component evaluations used here. A configuration missing even one benchmark cannot be scored the way the others are. Imputing the gap would mean inventing a number, and averaging over what happens to be present would quietly reward models for the benchmarks they skipped.
The rule is mechanical on purpose. Anyone can run it against the leaderboard and get the same list, which is what makes 'what is missing' an answerable question rather than a matter of who we happened to think of.
In practice it is a Terminal-Bench rule. In practice the coverage rule is a Terminal-Bench v2.1 rule. It is the newest of the seven and Artificial Analysis has run it on 229 of 637 records, so it is the single missing evaluation for 395 of the 434 models that cannot be scored here. Whole labs are excluded by it alone, Amazon's Nova line and Microsoft's Phi among them. That is a property of benchmark coverage, not of those models.
This is not the Intelligence Index. The Artificial Analysis Intelligence Index weights nine evaluations. The two omitted here, GDPval-AA at 20% and tau3-Banking at 14%, are long-horizon agentic work and out of scope by design. They are also the expensive ones: GDPval alone can be around 59% of a model's published cost per index task. Cost per solve here is therefore a different and much smaller number than the cost per task Artificial Analysis publishes, and the two should never be quoted against each other.
Muse Spark 1.3 at max effort. Scored on all seven evaluations and would rank near the top, but Artificial Analysis publishes no price for it at all and lists no serving host. A configuration with no price cannot enter a cost ranking. Its xhigh sibling is priced and does appear.
Four non-reasoning variants of Kimi and DeepSeek. Missing Terminal-Bench v2.1 entirely, and their index is flagged estimated rather than measured.
Five vendor-deprecated configurations. Superseded by a newer release. This is a buying guide, so a model you cannot adopt going forward is out, though it is named here rather than quietly dropped.
2
Benchmarks
Seven component evaluations of the Artificial Analysis Intelligence Index v4.1.1. Six are scored for accuracy and enter cost per solve; the seventh, AA-Omniscience, is the honesty gate and is never averaged in.
Table 1. Benchmarks in scope, what each measures, and how much it separates the field.
Benchmark
What it measures, and why it is in scope
GPQA Diamond
Graduate-level science, multiple choice. The floor check. Nearly every current model clears it, so it separates almost nothing.
saturated
HLE
Humanity's Last Exam: hard closed-ended questions across many fields. Still separates tiers sharply. One of the three that carry real signal.
discriminates
CritPt
Research-level physics reasoning. The hardest test in scope. Two cheap configurations score exactly zero while still billing.
discriminates
SciCode
Scientific coding, graded on sub-problems. Compressed. Ranks three through forty span nine points.
saturated
Terminal-Bench v2.1
Short agentic coding in a terminal, minutes per task. The one borderline inclusion: agentic, but bounded. Removing it shifts costs about 15% and does not change the ranking.
discriminates
AA-LCR
Reasoning over documents around 100,000 tokens. Compressed. The cheap tier is genuinely competitive here.
saturated
AA-Omniscience
Knowledge, with a penalty for confident errors. Not scored for accuracy. It is the honesty gate.
the gate
3
Metrics
3.1
Cost per solve
Take the price of one attempt on a benchmark and divide it by the share of problems the configuration got right. That is what a correct answer costs when you can spot a miss and retry. Do this on each of the six accuracy benchmarks, then combine them with a geometric mean so no single hard or expensive benchmark dominates.
cost per solvee = cost per attempte ÷ share solvede index cost per solve = geometric mean over the six accuracy evaluations e
Worked example
GPT-5.6 Sol at medium solves 92.6% of GPQA Diamond and an attempt costs $0.0157, so one correct answer costs $0.0157 divided by 0.926, which is $0.017. Repeat across the other five benchmarks and take the geometric mean: $0.14 per solve. Claude Sonnet 5 at max costs less per token but solves fewer problems per attempt, so it lands at $1.33.
3.2
Efficacy
On each of the six accuracy benchmarks, divide this configuration's score by the highest score any configuration reached on that benchmark. Average those six shares. 100% means it matched the leader everywhere.
3.3
Honesty and the noise band
On the AA-Omniscience knowledge test a correct answer scores plus one, a confident wrong answer minus one, and declining to answer zero. The total is expressed per 100 questions, so the scale runs from minus 100 to plus 100. Passing means above zero by more than the noise band described next.
AA-Omniscience is 6,000 questions, each scoring plus one, minus one or zero. The net is a mean of bounded scores, so its standard error is at most about 1.3 points on the published per-100 scale. A configuration within two standard errors of zero is reported as too close to call rather than passed or failed, because the test cannot separate it from chance. Six configurations sit in that band, and one of them, GPT-5.6 Terra at max effort, nets plus 0.05, which is four hundredths of a standard error from zero.
3.4
The 85% bar
A stated target, not a fitted one. An elbow-detection rule was tried first and rejected: with four to eight frontier points it is decided by whichever two neighbours happen to tie.
4
Results
Of 71 configurations, 21 clear both gates and are ranked. GPT-5.6 Sol at medium ranks first at $0.14 per solve with 86% efficacy, followed by GPT-6 Astra at low at $0.17 and GPT-5.6 Sol at high at $0.19. The qualifying set spans 5 labs (Anthropic, Google, Meta, Moonshot and OpenAI) and a 15-fold cost range, from $0.14 to $2.04, for 12 points of efficacy (Figure 1).
Of the rest, 20 pass the honesty gate but fall under the 85% bar, 6 sit inside the noise band on honesty and are not ranked, 20 are confidently wrong more often than right, and 4 score zero on at least one benchmark, which leaves cost per solve undefined. The cheapest configuration measured, GPT-5.6 Luna at low at $0.013 per solve, is in the failing tier; without the gate it would rank first.
Figure 1. The 21 qualifying configurations by cost per solve; labs outside Anthropic, OpenAI and Google highlighted.
5
Sensitivity
Two judgement calls set who appears: the 85% bar and the benchmark set. Both move the answer, so both are published rather than footnoted.
Table 2. Configurations and labs that qualify as the efficacy bar moves; the honesty gate is held fixed.
One benchmark carries most of that. Remove CritPt and 13 more configurations clear the same bar, from 8 labs instead of 5. Research-level physics is where these models separate, and whether it belongs in your definition of everyday work is a judgement you should make rather than inherit.
6
Limitations
Data collected from benchmarks as of 1 September 2026. GPT-6 Astra was released 3 September 2026 and added from pages retrieved 4 September 2026. There are new releases and re-grades between index patches, so figures move. Zaun Research will update this index regularly.
The retry assumption. Cost per solve assumes a failed attempt can be detected and retried. That holds on graded, bounded work, which is the scope here. It does not hold on long agentic runs, which is one reason those are out of scope.
Per-token price is not in the ranking. It does not predict cost per solve. One model here lists at 40% of another per token and costs more per task, because it uses more tokens and more turns.
Effort settings are not comparable across models. One vendor’s medium is not another’s. Read a configuration as a whole, never the effort label on its own. GPT-5.6 Sol at medium and Claude Opus 5 at medium list within 20% of each other per output token ($20 against $25 per million), yet one CritPt attempt costs $0.13 on the first and $1.40 on the second. The gap is tokens spent per task, not price per token.
Open-weight rows are priced at the creator’s rate, which is the expensive end. Marked open. Third-party hosts frequently undercut it, in one case by three to four times, so those configurations are likely cheaper in practice than shown here. The two NVIDIA rows are different again: NVIDIA offers no first-party API, so they carry a cross-provider median and are marked median price.
Platform premiums are not applied. Prices are first-party. Buying the same model through a cloud marketplace can add materially to it.
Long-running work is out of scope. Every benchmark here is bounded. The escalation triggers above point at the tail; this index does not rank models for it.
7
Data and reproducibility
Every figure on this page is computed from a dated snapshot of public sources, and the scripts that turn the snapshot into the ranking are versioned with it. Per-evaluation scores and cost per attempt were read from the Artificial Analysis evaluation pages [3–9] and checked against the published cost per task on each; index composition and weights from the Artificial Analysis methodology [2]; model-level cost per index task and per-token prices from the Artificial Analysis model pages [1]. First-party per-token prices for Anthropic and OpenAI were confirmed against the vendors’ own pricing pages [10, 11]; the cloud-platform multipliers discussed in the limitations come from the Bedrock, Azure and Vertex price lists [12–14]. Every other vendor’s first-party price is as recorded by Artificial Analysis on that model’s page [1]. The task-length figures are Zaun internal telemetry [15], August 2026, one prompt measured from submission to its last completion event, published in aggregate only. This is version 0.1 of the index, a preliminary snapshot; later versions will cite the snapshot date they supersede.
References
Artificial Analysis, model pages and Intelligence Index leaderboard, v4.1.1. artificialanalysis.ai/models. Retrieved 1 September 2026; GPT-6 Astra pages 4 September 2026.