ZaunZaun

Zaun Research

The Everyday Model IndexThe cheapest model you can trust by default.

Which model configuration should your AI agents run by default? We call that configuration the everyday model, and this index picks it with public benchmarks and one costing method.

The decision

What should be your default model?

The everyday model is the configuration that handles the nine in ten prompts that finish quickly. The top tier is held back for tasks that are genuinely hard or long.

Why it matters

9 in 10 tasks do not need the top-tier model.

94% of coding prompts and 92% of knowledge-work prompts finish inside 15 minutes and carry 71% and 50% of token spend. On those prompts a correct answer costs $0.14 on the cheapest qualifier against $2.04 on Claude Fable 5.1 at max, a 15× difference at 86% efficacy against 98%. That is the saving the default decides.

How we judged

Honest first, capable second, then cheapest.

A confident wrong answer costs labor hours, so honesty is a gate. Then at least 85% efficacy on a broad set of PhD-level, rigorous evaluations. Then rank by cost per correct answer.

Most work is short

Across our own agent telemetry, more than nine in ten prompts resolve inside fifteen minutes. These are the mundane and moderately complex tasks that make up a working day. Run them on the everyday model, the cheapest configuration you can trust, and hold the top tier for the rest.

94%
of coding-agent prompts resolve in under 15 minutes
17,609 measured prompts
92%
of knowledge-work prompts resolve in under 15 minutes
3,796 measured prompts
Coding agents41%33%20%94%Knowledge work42%31%19%92%under 15 minunder 1 min1 to 4 min4 to 16 min16 to 64 minover 64 min
Where the saving is, and where it is not

Those sub-15-minute prompts carry 71% of coding spend and 50% of knowledge-work spend, and that is where the everyday model pays: a correct answer costs $0.14 on the cheapest qualifier against $2.04 on Claude Fable 5.1 at max, a 15× difference at 86% efficacy against 98%. Start every prompt there and the majority of the bill moves to the cheap end.

The caveat is the tail. It is small by count and large by cost, larger in knowledge work than in coding, so an everyday model is not a licence to stop paying attention to it. The escalation rule below handles it. Route on predicted duration, not on prompt count, or the saving will disappoint.

Zaun internal data · August 2026 · a prompt is measured from submission to its last completion event

Three questions

Every configuration is asked the same three, in this order. Fail the first and cost never matters; fail the second and it does not reach the ranking.

First

Is it honest?

Right answers minus confidently wrong ones. Below zero it asserts more than it knows. Confident errors are paid for in labor hours, not tokens, so a model that asserts more than it knows is out before cost is even considered.

Second

Is it capable enough?

How close it gets to the best score on each test, averaged across six benchmarks. The six are graduate and research-level science, frontier exams, scientific coding, terminal work and long-document reasoning. A configuration that holds 85% of the best score across them is more than capable enough for everyday work; that is the bar to enter the ranking.

Third

What does it cost?

What one correct answer costs once you have paid for the attempts that failed. This is what the ranking sorts on.

The ranking

Clear the honesty gate by more than measurement noise, keep at least 85% efficacy, then rank by cost per solve. Cheapest first. A configuration is a model plus a reasoning-effort setting. GPT-5.6 Sol at medium and GPT-5.6 Sol at max are two configurations of one model, and they rank differently. Open any row for its per-benchmark detail and its weak spot. Every configuration measured is here, including the ones that do not qualify and why.

#ConfigurationCost per solveCost per solveWhat one correct answer costs once you have paid for the attempts that failed.EfficacyEfficacyHow close it gets to the best score on each test, averaged across six benchmarks.HonestyHonestyRight answers minus confidently wrong ones. Below zero it asserts more than it knows.
1GPT-5.6 Solmedium$0.1486%passes
Weak spotCritPt at 71%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+19.4per 100 questions, right answers minus confident errors
Cost per fact$0.05What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA92.6%96%$0.0157$0.017
HLE42.2%71%$0.0430$0.102
CritPt22.9%71%$0.1259$0.550
SciCode56.5%91%$0.0285$0.050
Terminal-Bench86.1%94%$0.2302$0.267
Long context74.3%89%$0.3830$0.515

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

2GPT-6 Astralow$0.1788%passes
Weak spotCritPt at 81%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+40.5per 100 questions, right answers minus confident errors
Cost per fact$0.02What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.1%97%$0.0155$0.017
HLE49.2%83%$0.0389$0.079
CritPt26.3%81%$0.1449$0.551
SciCode51.0%82%$0.0425$0.083
Terminal-Bench88.0%96%$0.2903$0.330
Long context72.7%87%$0.9501$1.308

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

3GPT-5.6 Solhigh$0.1989%passes
Weak spotHLE at 78%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+20.4per 100 questions, right answers minus confident errors
Cost per fact$0.09What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA92.8%96%$0.0243$0.026
HLE46.0%78%$0.0828$0.180
CritPt25.7%80%$0.2453$0.955
SciCode56.9%92%$0.0353$0.062
Terminal-Bench87.3%96%$0.2794$0.320
Long context75.3%90%$0.3859$0.512

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

4Muse Spark 1.3xhigh$0.2190%passes
Weak spotHLE at 80%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+23.1per 100 questions, right answers minus confident errors
Cost per fact$0.07What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA94.1%98%$0.0298$0.032
HLE47.5%80%$0.1041$0.219
CritPt26.0%81%$0.3080$1.185
SciCode58.6%95%$0.0371$0.063
Terminal-Bench85.4%93%$0.9380$1.099
Long context79.3%95%$0.1315$0.166

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

5Gemini 3.8 Flashhigh$0.2286%passes
Weak spotCritPt at 57%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+29.6per 100 questions, right answers minus confident errors
Cost per fact$0.04What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA95.3%99%$0.0245$0.026
HLE47.8%81%$0.0920$0.192
CritPt18.3%57%$0.2800$1.531
SciCode53.6%86%$0.0769$0.143
Terminal-Bench87.6%96%$0.7178$0.819
Long context81.0%97%$0.0988$0.122

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

6GPT-6 Astramedium$0.2691%passes
Weak spotSciCode at 82%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+42.2per 100 questions, right answers minus confident errors
Cost per fact$0.04What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.9%98%$0.0244$0.026
HLE52.7%89%$0.1005$0.191
CritPt29.1%90%$0.2851$0.978
SciCode51.0%82%$0.0509$0.100
Terminal-Bench89.5%98%$0.4506$0.503
Long context76.3%92%$0.9533$1.249

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

7GPT-5.6 Solxhigh$0.2691%passes
Weak spotHLE at 80%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+21.0per 100 questions, right answers minus confident errors
Cost per fact$0.16What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.1%97%$0.0370$0.040
HLE47.3%80%$0.1421$0.300
CritPt28.6%89%$0.4367$1.527
SciCode56.0%90%$0.0449$0.080
Terminal-Bench89.5%98%$0.3821$0.427
Long context76.3%92%$0.3876$0.508

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

8GPT-6 Astrahigh$0.3692%passes
Weak spotSciCode at 83%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+43.7per 100 questions, right answers minus confident errors
Cost per fact$0.06What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA94.9%99%$0.0371$0.039
HLE53.1%90%$0.1618$0.305
CritPt28.9%89%$0.4335$1.502
SciCode51.6%83%$0.0696$0.135
Terminal-Bench89.9%98%$0.6450$0.718
Long context76.0%91%$0.9597$1.263

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

9Claude Opus 5medium$0.3889%passes
Weak spotSciCode at 82%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+31.0per 100 questions, right answers minus confident errors
Cost per fact$0.02What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA91.9%96%$0.0241$0.026
HLE51.3%87%$0.1870$0.364
CritPt26.9%83%$1.4033$5.217
SciCode50.7%82%$0.0369$0.073
Terminal-Bench86.1%94%$0.7522$0.874
Long context78.7%94%$0.7378$0.938

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

10Claude Fable 5.1low$0.4090%passes
Weak spotHLE at 83%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+34.1per 100 questions, right answers minus confident errors
Cost per fact$0.03What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA88.1%92%$0.0164$0.019
HLE48.9%83%$0.1238$0.253
CritPt27.7%86%$1.2559$4.534
SciCode55.7%90%$0.0690$0.124
Terminal-Bench85.0%93%$0.6832$0.804
Long context79.3%95%$1.4690$1.853

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

11 more configurations qualify
11GPT-5.6 Solmax$0.4194%passes
Weak spotHLE at 84%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+22.0per 100 questions, right answers minus confident errors
Cost per fact$0.39What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA94.1%98%$0.0732$0.078
HLE49.5%84%$0.2684$0.542
CritPt32.3%100%$0.8591$2.660
SciCode56.1%91%$0.0797$0.142
Terminal-Bench88.0%96%$0.5435$0.618
Long context77.7%93%$0.4016$0.517

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

12Kimi K3maxopen$0.5089%passes
Weak spotCritPt at 72%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+19.7per 100 questions, right answers minus confident errors
Cost per fact$0.67What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.5%97%$0.1512$0.162
HLE46.9%79%$0.3936$0.839
CritPt23.4%72%$0.9019$3.854
SciCode58.7%95%$0.0602$0.103
Terminal-Bench85.0%93%$0.6271$0.738
Long context82.7%99%$0.3086$0.373

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

13Claude Fable 5.1medium$0.5192%passes
Weak spotSciCode at 89%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+37.6per 100 questions, right answers minus confident errors
Cost per fact$0.04What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA88.6%92%$0.0263$0.030
HLE53.8%91%$0.2254$0.419
CritPt29.1%90%$1.8039$6.199
SciCode55.3%89%$0.0756$0.137
Terminal-Bench88.0%96%$0.7691$0.874
Long context78.7%94%$1.4752$1.875

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

14Claude Opus 5high$0.5592%passes
Weak spotSciCode at 88%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+33.7per 100 questions, right answers minus confident errors
Cost per fact$0.03What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.7%97%$0.0452$0.048
HLE52.8%89%$0.3521$0.667
CritPt28.3%88%$1.9945$7.048
SciCode54.3%88%$0.0481$0.089
Terminal-Bench87.6%96%$1.2258$1.399
Long context76.3%92%$0.7456$0.977

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

15GPT-6 Astraxhigh$0.5593%passes
Weak spotSciCode at 80%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+43.4per 100 questions, right answers minus confident errors
Cost per fact$0.12What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA96.3%100%$0.0726$0.075
HLE54.6%92%$0.2532$0.464
CritPt31.4%97%$0.7980$2.539
SciCode49.5%80%$0.1216$0.245
Terminal-Bench89.1%98%$0.8704$0.977
Long context74.0%89%$0.9699$1.311

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

16Claude Fable 5.1high$0.7394%passes
Weak spotLong context at 92%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+40.8per 100 questions, right answers minus confident errors
Cost per fact$0.05What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA90.6%94%$0.0553$0.061
HLE55.9%95%$0.3970$0.710
CritPt30.3%94%$2.7477$9.068
SciCode57.6%93%$0.0891$0.155
Terminal-Bench89.9%98%$1.1876$1.321
Long context77.0%92%$1.4810$1.923

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

17Claude Opus 5xhigh$0.7392%passes
Weak spotCritPt at 86%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+35.4per 100 questions, right answers minus confident errors
Cost per fact$0.04What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.7%97%$0.0635$0.068
HLE54.4%92%$0.5171$0.951
CritPt27.7%86%$2.4581$8.874
SciCode55.0%89%$0.0723$0.132
Terminal-Bench88.0%96%$1.8677$2.122
Long context76.3%92%$0.7529$0.987

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

18GPT-6 Astramax$0.7694%passes
Weak spotSciCode at 87%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+43.4per 100 questions, right answers minus confident errors
Cost per fact$0.22What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA96.1%100%$0.0924$0.096
HLE54.7%93%$0.3781$0.691
CritPt31.7%98%$1.1662$3.677
SciCode54.1%87%$0.2323$0.430
Terminal-Bench88.4%97%$1.1997$1.357
Long context74.3%89%$0.9973$1.342

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

19Claude Opus 5max$0.9493%passes
Weak spotSciCode at 90%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+37.1per 100 questions, right answers minus confident errors
Cost per fact$0.07What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.2%97%$0.0921$0.099
HLE54.9%93%$0.6736$1.227
CritPt29.1%90%$2.8845$9.912
SciCode55.7%90%$0.1177$0.211
Terminal-Bench89.1%98%$2.4213$2.717
Long context75.7%91%$0.7611$1.005

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

20Claude Fable 5.1xhigh$1.3297%passes
Weak spotLong context at 94%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+42.4per 100 questions, right answers minus confident errors
Cost per fact$0.18What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.4%97%$0.0939$0.101
HLE58.7%99%$1.0254$1.747
CritPt31.1%96%$4.5338$14.578
SciCode60.1%97%$0.2410$0.401
Terminal-Bench91.0%100%$2.4726$2.717
Long context78.0%94%$1.5078$1.933

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

21Claude Fable 5.1max$2.0498%passes
Weak spotCritPt at 92%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+43.5per 100 questions, right answers minus confident errors
Cost per fact$0.62What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.7%97%$0.1378$0.147
HLE59.1%100%$1.5859$2.683
CritPt29.7%92%$5.7138$19.238
SciCode62.0%100%$0.6585$1.062
Terminal-Bench91.4%100%$4.1512$4.542
Long context80.0%96%$1.5845$1.981

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

20 are honest but fall under the 85% efficacy bar

20 of these 20 are held back by the same benchmark: CritPt, research-level physics. Most clear the other five comfortably. Gemini 3.7 Flash at high (84%), Muse Spark 1.2 at xhigh (84%), Grok 4.6 at xhigh (84%) sit within two points of the bar, so a single benchmark is deciding whether they qualify.

·GLM-5.3-Flashdefaultopen78% efficacy$0.04278%passes
Weak spotCritPt at 48%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+7.5per 100 questions, right answers minus confident errors
Cost per fact$0.0080What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA91.2%95%$0.0055$0.006
HLE39.9%67%$0.0217$0.054
CritPt15.4%48%$0.0420$0.272
SciCode46.1%74%$0.0140$0.030
Terminal-Bench84.3%92%$0.0762$0.090
Long context78.0%94%$0.0169$0.022

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Gemini 3.7 Flashlow73% efficacy$0.04673%passes
Weak spotCritPt at 18%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+22.1per 100 questions, right answers minus confident errors
Cost per fact$0.0050What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA90.1%94%$0.0036$0.004
HLE35.1%59%$0.0059$0.017
CritPt5.7%18%$0.0230$0.404
SciCode53.6%87%$0.0045$0.008
Terminal-Bench79.8%87%$0.3184$0.399
Long context78.3%94%$0.0836$0.107

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Gemini 3.8 Flashlow74% efficacy$0.05874%passes
Weak spotCritPt at 12%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+21.4per 100 questions, right answers minus confident errors
Cost per fact$0.0039What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA92.0%96%$0.0042$0.004
HLE37.1%63%$0.0083$0.022
CritPt4.0%12%$0.0295$0.738
SciCode54.3%88%$0.0043$0.008
Terminal-Bench83.2%91%$0.4927$0.593
Long context78.7%94%$0.0837$0.106

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Gemini 3.7 Flashmedium78% efficacy$0.06178%passes
Weak spotCritPt at 29%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+23.7per 100 questions, right answers minus confident errors
Cost per fact$0.0093What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA92.1%96%$0.0049$0.005
HLE39.0%66%$0.0097$0.025
CritPt9.4%29%$0.0392$0.417
SciCode57.9%93%$0.0100$0.017
Terminal-Bench78.3%86%$0.4048$0.517
Long context81.0%97%$0.0856$0.106

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Grok 4.6low68% efficacy$0.08968%passes
Weak spotCritPt at 18%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+25.9per 100 questions, right answers minus confident errors
Cost per fact$0.01What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA87.9%91%$0.0075$0.009
HLE27.6%47%$0.0132$0.048
CritPt5.7%18%$0.0391$0.686
SciCode48.4%78%$0.0098$0.020
Terminal-Bench75.3%82%$0.2647$0.351
Long context78.7%94%$0.1962$0.249

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Gemini 3.8 Flashmedium81% efficacy$0.1181%passes
Weak spotCritPt at 38%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+28.6per 100 questions, right answers minus confident errors
Cost per fact$0.01What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.5%97%$0.0070$0.007
HLE42.1%71%$0.0248$0.059
CritPt12.3%38%$0.1017$0.827
SciCode54.4%88%$0.0294$0.054
Terminal-Bench83.9%92%$0.5419$0.646
Long context82.0%98%$0.0897$0.109

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Sollow78% efficacy$0.1178%passes
Weak spotCritPt at 46%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+18.9per 100 questions, right answers minus confident errors
Cost per fact$0.03What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA89.8%93%$0.0101$0.011
HLE39.4%67%$0.0224$0.057
CritPt14.9%46%$0.0751$0.504
SciCode55.4%89%$0.0241$0.043
Terminal-Bench76.8%84%$0.1665$0.217
Long context73.0%88%$0.3817$0.523

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Gemini 3.7 Flashhigh84% efficacy$0.1184%passes
Weak spotCritPt at 44%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+26.5per 100 questions, right answers minus confident errors
Cost per fact$0.01What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA94.5%98%$0.0100$0.011
HLE47.9%81%$0.0368$0.077
CritPt14.3%44%$0.1148$0.803
SciCode56.8%92%$0.0245$0.043
Terminal-Bench85.8%94%$0.4940$0.576
Long context80.0%96%$0.0893$0.112

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Kimi K3lowopen67% efficacy$0.1567%passes
Weak spotCritPt at 10%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+3.9per 100 questions, right answers minus confident errors
Cost per fact$0.13What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA84.2%88%$0.0123$0.015
HLE25.0%42%$0.0242$0.097
CritPt3.1%10%$0.0503$1.599
SciCode51.2%83%$0.0184$0.036
Terminal-Bench82.4%90%$0.3107$0.377
Long context77.0%92%$0.2931$0.381

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-6 Astranon-reasoning80% efficacy$0.1780%passes
Weak spotCritPt at 61%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+26.6per 100 questions, right answers minus confident errors
Cost per fact$0.0077What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA89.5%93%$0.0110$0.012
HLE37.1%63%$0.0190$0.051
CritPt19.7%61%$0.1328$0.674
SciCode50.5%81%$0.0463$0.092
Terminal-Bench89.1%98%$0.4432$0.497
Long context68.0%82%$0.9476$1.393

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Grok 4.6medium82% efficacy$0.2282%passes
Weak spotCritPt at 55%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+28.0per 100 questions, right answers minus confident errors
Cost per fact$0.02What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.5%97%$0.0302$0.032
HLE42.1%71%$0.0937$0.223
CritPt17.7%55%$0.2419$1.367
SciCode54.6%88%$0.0233$0.043
Terminal-Bench84.3%92%$0.8093$0.960
Long context72.7%87%$0.2021$0.278

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Claude Opus 5low82% efficacy$0.2382%passes
Weak spotCritPt at 72%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+28.6per 100 questions, right answers minus confident errors
Cost per fact$0.01What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA88.9%92%$0.0076$0.009
HLE43.4%73%$0.0625$0.144
CritPt23.1%72%$0.7204$3.119
SciCode48.0%77%$0.0287$0.060
Terminal-Bench76.4%84%$0.4984$0.652
Long context77.0%92%$0.7306$0.949

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Grok 4.6high83% efficacy$0.2583%passes
Weak spotCritPt at 53%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+30.5per 100 questions, right answers minus confident errors
Cost per fact$0.03What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA94.9%99%$0.0364$0.038
HLE42.9%73%$0.1205$0.281
CritPt17.1%53%$0.3160$1.848
SciCode53.6%87%$0.0272$0.051
Terminal-Bench88.4%97%$0.8098$0.916
Long context75.0%90%$0.2026$0.270

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Muse Spark 1.2xhigh84% efficacy$0.2684%passes
Weak spotCritPt at 55%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+27.2per 100 questions, right answers minus confident errors
Cost per fact$0.04What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA90.4%94%$0.0394$0.044
HLE45.5%77%$0.1006$0.221
CritPt17.7%55%$0.3458$1.952
SciCode56.4%91%$0.0605$0.107
Terminal-Bench80.1%88%$0.7972$0.995
Long context83.3%100%$0.1332$0.160

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Grok 4.6xhigh84% efficacy$0.2784%passes
Weak spotCritPt at 61%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+29.3per 100 questions, right answers minus confident errors
Cost per fact$0.03What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.5%97%$0.0383$0.041
HLE44.1%75%$0.1248$0.283
CritPt19.7%61%$0.3109$1.578
SciCode51.6%83%$0.0296$0.057
Terminal-Bench88.0%96%$1.1621$1.321
Long context75.7%91%$0.2036$0.269

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GLM-5.3maxopen83% efficacy$0.3283%passes
Weak spotCritPt at 59%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+14.3per 100 questions, right answers minus confident errors
Cost per fact$0.08What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA91.7%95%$0.0316$0.035
HLE42.3%72%$0.2328$0.550
CritPt19.1%59%$0.4522$2.368
SciCode56.5%91%$0.0820$0.145
Terminal-Bench83.9%92%$0.6583$0.785
Long context76.3%92%$0.1486$0.195

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Gemini 3.1 Pro Previewdefault84% efficacy$0.3784%passes
Weak spotCritPt at 55%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+31.9per 100 questions, right answers minus confident errors
Cost per fact$0.06What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA94.1%98%$0.0715$0.076
HLE47.0%80%$0.1968$0.419
CritPt17.7%55%$0.4739$2.677
SciCode58.9%95%$0.0887$0.151
Terminal-Bench73.8%81%$0.5780$0.783
Long context79.0%95%$0.2127$0.269

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Qwen3.8 2.4T A95Bdefaultopen82% efficacy$0.4082%passes
Weak spotCritPt at 62%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+4.3per 100 questions, right answers minus confident errors
Cost per fact$0.65What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA93.5%97%$0.0658$0.070
HLE42.4%72%$0.2016$0.475
CritPt20.0%62%$0.4828$2.414
SciCode51.6%83%$0.1006$0.195
Terminal-Bench82.0%90%$0.6941$0.846
Long context75.3%90%$0.2220$0.295

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Qwen3.8 Maxdefault82% efficacy$0.4582%passes
Weak spotCritPt at 62%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+3.4per 100 questions, right answers minus confident errors
Cost per fact$0.84What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA92.7%96%$0.0735$0.079
HLE43.0%73%$0.2234$0.519
CritPt20.0%62%$0.5092$2.546
SciCode52.9%85%$0.1106$0.209
Terminal-Bench81.3%89%$1.0157$1.250
Long context74.3%89%$0.2218$0.298

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Claude Sonnet 5max81% efficacy$1.3381%passes
Weak spotCritPt at 52%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+16.4per 100 questions, right answers minus confident errors
Cost per fact$0.18What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA91.1%95%$0.3663$0.402
HLE41.3%70%$1.0208$2.472
CritPt16.9%52%$2.3138$13.691
SciCode53.6%87%$0.1917$0.358
Terminal-Bench80.5%88%$1.8879$2.345
Long context77.0%92%$0.3679$0.478

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

6 are too close to call on honesty

Their net knowledge score sits within measurement noise of zero, so the test cannot say whether they are right more often than confidently wrong. They are not ranked, because a cost per correct answer needs a denominator we can stand behind. The closest case is GPT-5.6 Terra at max, which nets +0.05 against a standard error of about 1.29, so it is 0.04 standard errors from zero.

·GPT-5.6 Solnon-reasoninghonesty +1.10$0.1159%unclear
Weak spotCritPt at 16%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+1.1per 100 questions, right answers minus confident errors
Cost per fact$0.05What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA79.0%82%$0.0036$0.005
HLE16.7%28%$0.0055$0.033
CritPt5.1%16%$0.0471$0.923
SciCode47.1%76%$0.0189$0.040
Terminal-Bench74.2%81%$0.2871$0.387
Long context57.0%68%$0.3789$0.665

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·MiniMax-M3defaultopenhonesty +1.35$0.1169%unclear
Weak spotCritPt at 12%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+1.4per 100 questions, right answers minus confident errors
Cost per fact$0.26What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA92.9%97%$0.0099$0.011
HLE39.0%66%$0.0259$0.067
CritPt3.7%12%$0.0955$2.572
SciCode45.4%73%$0.0259$0.057
Terminal-Bench65.2%71%$0.2814$0.432
Long context80.3%96%$0.0298$0.037

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Claude Sonnet 5non-reasoninghonesty −0.65$0.2060%unclear
Weak spotCritPt at 3%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−0.7per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA80.0%83%$0.0087$0.011
HLE19.0%32%$0.0201$0.106
CritPt1.1%3%$0.0733$6.664
SciCode48.6%78%$0.0131$0.027
Terminal-Bench75.3%82%$0.8877$1.179
Long context66.0%79%$0.1921$0.291

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·DeepSeek V4 Pro 0813maxopenhonesty +0.83$0.2780%unclear
Weak spotCritPt at 56%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+0.8per 100 questions, right answers minus confident errors
Cost per fact$1.51What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA92.8%96%$0.0598$0.064
HLE41.0%69%$0.1357$0.331
CritPt18.0%56%$0.3568$1.982
SciCode49.2%79%$0.0503$0.102
Terminal-Bench78.7%86%$0.4124$0.524
Long context75.3%90%$0.1389$0.184

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Nemotron 3 Ultra 550B A55Breasoningmedian pricehonesty −0.40$0.3159%unclear
Weak spotCritPt at 10%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−0.4per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA86.7%90%$0.0396$0.046
HLE28.4%48%$0.0996$0.351
CritPt3.1%10%$0.2226$7.082
SciCode39.9%64%$0.0124$0.031
Terminal-Bench53.9%59%$1.1414$2.116
Long context71.0%85%$0.0845$0.119

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Terramaxhonesty +0.05$0.3890%unclear
Weak spotHLE at 73%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+0.1per 100 questions, right answers minus confident errors
Cost per fact$140.60What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA92.5%96%$0.0730$0.079
HLE42.9%73%$0.2277$0.531
CritPt30.0%93%$0.6084$2.028
SciCode53.9%87%$0.1144$0.212
Terminal-Bench88.0%96%$0.5732$0.651
Long context79.7%96%$0.2130$0.267

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

20 fail the honesty gate at every effort setting

These are right less often than they are confidently wrong. Several are also the cheapest configurations measured: GPT-5.6 Luna (low) would top the ranking at $0.013 per solve, roughly 10x below the cheapest qualifier. They are excluded anyway, because every confident error is undone by a person at a labor rate that swamps the saving. Note that they look fine on a standard test: their median GPQA Diamond score is 90%.

·GPT-5.6 Lunalowhonesty −15$0.01355%fails
Weak spotCritPt at 8%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−14.7per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA83.5%87%$0.0008$0.001
HLE19.8%34%$0.0014$0.007
CritPt2.6%8%$0.0037$0.142
SciCode45.6%74%$0.0015$0.003
Terminal-Bench43.4%48%$0.0239$0.055
Long context65.3%78%$0.0191$0.029

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Lunamediumhonesty −13$0.01561%fails
Weak spotCritPt at 15%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−13.2per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA85.9%89%$0.0013$0.002
HLE25.8%44%$0.0029$0.011
CritPt4.9%15%$0.0066$0.135
SciCode45.8%74%$0.0018$0.004
Terminal-Bench53.2%58%$0.0243$0.046
Long context72.0%86%$0.0193$0.027

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Lunanon-reasoninghonesty −25$0.01539%fails
Weak spotCritPt at 1%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−24.9per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA64.5%67%$0.0002$0.000
HLE7.2%12%$0.0003$0.004
CritPt0.3%1%$0.0023$0.767
SciCode39.9%64%$0.0010$0.003
Terminal-Bench39.0%43%$0.0405$0.104
Long context38.7%46%$0.0190$0.049

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Lunahighhonesty −12$0.02175%fails
Weak spotCritPt at 51%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−12.0per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA89.2%93%$0.0034$0.004
HLE33.4%56%$0.0095$0.028
CritPt16.6%51%$0.0190$0.115
SciCode50.7%82%$0.0029$0.006
Terminal-Bench69.7%76%$0.0321$0.046
Long context74.0%89%$0.0196$0.026

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Lunaxhighhonesty −11$0.03079%fails
Weak spotHLE at 63%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−10.8per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA89.5%93%$0.0059$0.007
HLE37.0%63%$0.0167$0.045
CritPt20.6%64%$0.0395$0.192
SciCode50.0%81%$0.0045$0.009
Terminal-Bench77.9%85%$0.0374$0.048
Long context73.3%88%$0.0199$0.027

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Lunamaxhonesty −10$0.04482%fails
Weak spotCritPt at 64%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−10.3per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA91.1%95%$0.0099$0.011
HLE39.5%67%$0.0293$0.074
CritPt20.6%64%$0.0633$0.307
SciCode52.5%85%$0.0093$0.018
Terminal-Bench80.9%89%$0.0529$0.065
Long context78.3%94%$0.0209$0.027

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Qwen3.8-Flash-Nextdefaultopenhonesty −10$0.05376%fails
Weak spotCritPt at 35%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−9.7per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA92.3%96%$0.0083$0.009
HLE38.0%64%$0.0240$0.063
CritPt11.1%35%$0.0557$0.500
SciCode46.9%76%$0.0126$0.027
Terminal-Bench86.1%94%$0.1060$0.123
Long context77.0%92%$0.0175$0.023

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·DeepSeek V4 Prohighopenhonesty −11$0.06268%fails
Weak spotCritPt at 31%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−10.6per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA90.5%94%$0.0092$0.010
HLE35.2%60%$0.0286$0.081
CritPt10.0%31%$0.0653$0.653
SciCode46.4%75%$0.0046$0.010
Terminal-Bench64.8%71%$0.1051$0.162
Long context67.0%80%$0.0431$0.064

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Terranon-reasoninghonesty −23$0.08450%fails
Weak spotCritPt at 6%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−23.2per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA74.6%78%$0.0022$0.003
HLE11.4%19%$0.0030$0.026
CritPt2.0%6%$0.0255$1.275
SciCode44.6%72%$0.0100$0.022
Terminal-Bench56.2%62%$0.2535$0.451
Long context55.0%66%$0.1896$0.345

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Terralowhonesty −7$0.08466%fails
Weak spotCritPt at 29%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−6.8per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA84.3%88%$0.0060$0.007
HLE29.2%49%$0.0142$0.049
CritPt9.4%29%$0.0434$0.462
SciCode49.2%79%$0.0119$0.024
Terminal-Bench62.5%68%$0.2033$0.325
Long context68.7%82%$0.1907$0.278

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·DeepSeek V4 Promaxopenhonesty −11$0.09172%fails
Weak spotCritPt at 40%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−10.7per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA88.8%92%$0.0252$0.028
HLE37.5%64%$0.0464$0.123
CritPt12.9%40%$0.0929$0.723
SciCode50.0%81%$0.0114$0.023
Terminal-Bench64.0%70%$0.0949$0.148
Long context70.0%84%$0.0464$0.066

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Terramediumhonesty −5$0.09474%fails
Weak spotCritPt at 54%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−5.1per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA87.2%91%$0.0093$0.011
HLE33.3%56%$0.0255$0.077
CritPt17.4%54%$0.0696$0.400
SciCode49.7%80%$0.0150$0.030
Terminal-Bench72.3%79%$0.1824$0.252
Long context70.3%84%$0.1914$0.272

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·DeepSeek V4 Flash 0731maxopenhonesty −14$0.1078%fails
Weak spotCritPt at 51%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−14.3per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA90.8%94%$0.0125$0.014
HLE38.6%65%$0.0429$0.111
CritPt16.6%51%$0.1145$0.691
SciCode49.9%81%$0.0422$0.085
Terminal-Bench78.7%86%$0.1407$0.179
Long context74.3%89%$0.0476$0.064

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·DeepSeek V4 Flash Visionmaxhonesty −18$0.1173%fails
Weak spotCritPt at 34%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−17.6per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA91.3%95%$0.0169$0.018
HLE34.5%58%$0.0379$0.110
CritPt10.9%34%$0.1105$1.018
SciCode46.6%75%$0.0344$0.074
Terminal-Bench74.2%81%$0.1079$0.145
Long context78.0%94%$0.0476$0.061

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Nemotron 3 Super 120B A12Breasoningmedian pricehonesty −42$0.1150%fails
Weak spotCritPt at 10%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−41.5per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA80.0%83%$0.0186$0.023
HLE20.8%35%$0.0290$0.140
CritPt3.1%10%$0.0485$1.542
SciCode36.0%58%$0.0022$0.006
Terminal-Bench38.6%42%$0.5551$1.439
Long context60.3%72%$0.0261$0.043

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Terrahighhonesty −3$0.1580%fails
Weak spotHLE at 65%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−3.5per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA89.6%93%$0.0184$0.021
HLE38.5%65%$0.0619$0.161
CritPt22.9%71%$0.1741$0.760
SciCode50.1%81%$0.0232$0.046
Terminal-Bench75.7%83%$0.2744$0.362
Long context73.3%88%$0.1929$0.263

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Qwen3.8 27Bnon-reasoningopenhonesty −8$0.1849%fails
Weak spotCritPt at 1%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−8.0per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA81.8%85%$0.0113$0.014
HLE12.1%21%$0.0128$0.105
CritPt0.3%1%$0.0324$11.356
SciCode35.6%57%$0.0048$0.013
Terminal-Bench49.1%54%$0.8171$1.665
Long context63.0%76%$0.0551$0.087

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·GPT-5.6 Terraxhighhonesty −3$0.1985%fails
Weak spotHLE at 71%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−3.0per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA90.8%94%$0.0257$0.028
HLE41.9%71%$0.0957$0.228
CritPt27.1%84%$0.2907$1.073
SciCode51.6%83%$0.0325$0.063
Terminal-Bench80.1%88%$0.3356$0.419
Long context75.0%90%$0.1948$0.260

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Kimi K2.7 Codedefaultopenhonesty −10$0.2371%fails
Weak spotCritPt at 31%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−10.2per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA89.6%93%$0.0396$0.044
HLE35.0%59%$0.1047$0.299
CritPt10.0%31%$0.2563$2.563
SciCode47.5%77%$0.0238$0.050
Terminal-Bench67.4%74%$0.4143$0.615
Long context75.0%90%$0.0969$0.129

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Qwen3.8 27Bxhighopenhonesty −10$0.3270%fails
Weak spotCritPt at 17%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−10.0per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA90.5%94%$0.0464$0.051
HLE33.9%57%$0.1326$0.391
CritPt5.4%17%$0.3217$5.926
SciCode44.7%72%$0.0645$0.144
Terminal-Bench79.8%87%$0.5963$0.748
Long context77.3%93%$0.0617$0.080

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

4 never solve one of the benchmarks at all
·Gemini 3.5 Flash-Litedefaultscores zero on CritPtno price56%passes
Weak spotCritPt at 0%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin+5.2per 100 questions, right answers minus confident errors
Cost per fact$0.06What one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA83.8%87%$0.0122$0.015
HLE18.8%32%$0.0271$0.144
CritPt0.0%0%$0.0704never solves
SciCode40.9%66%$0.0126$0.031
Terminal-Bench53.6%59%$0.2571$0.480
Long context74.7%90%$0.0390$0.052

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Claude Haiku 4.5reasoningscores zero on CritPtno price49%fails
Weak spotCritPt at 0%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−4.4per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA67.2%70%$0.0749$0.112
HLE10.4%18%$0.1245$1.197
CritPt0.0%0%$0.1648never solves
SciCode43.3%70%$0.0507$0.117
Terminal-Bench44.2%48%$0.6990$1.581
Long context73.7%88%$0.1213$0.165

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Qwen3.8 27Bmediumopenscores zero on CritPtno price56%fails
Weak spotCritPt at 0%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−36.1per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA84.5%88%$0.0152$0.018
HLE14.1%24%$0.0390$0.277
CritPt0.0%0%$0.0918never solves
SciCode38.1%61%$0.0117$0.031
Terminal-Bench65.2%71%$0.4084$0.627
Long context76.3%92%$0.0576$0.075

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

·Qwen3.8 27Blowopenscores zero on CritPtno price57%fails
Weak spotCritPt at 0%The benchmark where it falls furthest behind the best model, and how far.
Honesty margin−26.7per 100 questions, right answers minus confident errors
Cost per factnot definedWhat one trustworthy answer costs on the knowledge test, after confident mistakes cancel out correct ones.
BenchmarkScoreShare of bestPer attemptPer solve
GPQA84.5%88%$0.0141$0.017
HLE14.0%24%$0.0217$0.154
CritPt0.0%0%$0.0509never solves
SciCode39.8%64%$0.0082$0.021
Terminal-Bench67.4%74%$0.4303$0.638
Long context74.7%90%$0.0571$0.076

The six accuracy benchmarks. The seventh, AA-Omniscience, is the honesty gate shown above, and is never averaged into cost per solve.

Cheapest qualifying configuration $0.14 · most expensive $2.04 · a 15× range for 12 points of efficacy

What the data shows

Four findings, each with the point first and the evidence behind it one click down.

3.3×Use the reasoning lever, but keep the big model.

Claude Sonnet 5 at max effort costs less per token than Claude Fable 5.1, and 3.3 times more per correct answer, because it uses more tokens and solves fewer problems per attempt. Fable 5.1 at low effort is cheaper and 9 points more capable. The sticker price told the opposite story.

Claude Sonnet 5 max$1.3381% efficacyClaude Opus 5 medium$0.3889% efficacyClaude Fable 5.1 low$0.4090% efficacycost per solve, same vendor, all pass the honesty gate
Sonnet 5 max $1.33 per solve at 81% efficacy · Fable 5.1 low $0.40 at 90% · same vendor, both pass the honesty gate
20 of 71The cheapest configurations are among the ones that make things up.

20 configurations are right less often than they are confidently wrong. GPT-5.6 Luna (low) would top the whole ranking at $0.013 per solve if the gate did not exist. It is excluded because every confident error is undone by a person, at a labor rate that swamps the saving. These models look fine on a standard test: their median GPQA Diamond score is 90%. Honesty is not something you can read off accuracy.

noise band−40−20+0+20+40$0.03$0.1$0.3$1every GPT-5.6 Luna settingGPT-5.6 Sol medium, cheapest to qualifycost per solve, log scale · vertical: honesty margin per 100 questions
Green passes, amber is within measurement noise of zero, red fails · every GPT-5.6 Luna setting fails, best margin −10.3
The effort setting moves cost more than the model name does.

GPT-5.6 Sol alone spans $0.14 to $0.41 per solve and 86% to 94% efficacy across its effort settings. That is most of the useful range without changing model. The same is true of every frontier family here, which is why a configuration, model plus effort, is the unit ranked and never the model alone.

85% bar60%80%100%non-reasoninglowmediumhighxhighmax$0.11$0.14$0.19$0.26$0.41GPT-5.6 Sol, one model, six effort settings · cost per solve against efficacy
Sol at low (78%) falls under the bar; at non-reasoning its honesty is too close to call · the setting matters on both gates
1 of 10Labs outside the big three hold top-ten places.

Muse Spark 1.3 ranks 4. Of the 21 configurations that qualify, one has open weights and can be self-hosted. The field is wider than the three names most buying conversations start with, and it is moving fast: GPT-6 Astra, released 3 September 2026, was added on 4 September 2026 and holds 3 of the top 10 places.

1. GPT-5.6 Sol medium$0.142. GPT-6 Astra low$0.173. GPT-5.6 Sol high$0.194. Muse Spark 1.3 xhigh$0.215. Gemini 3.8 Flash high$0.226. GPT-6 Astra medium$0.267. GPT-5.6 Sol xhigh$0.268. GPT-6 Astra high$0.369. Claude Opus 5 medium$0.3810. Claude Fable 5.1 low$0.4011. GPT-5.6 Sol max$0.4112. Kimi K3 max$0.5013. Claude Fable 5.1 medium$0.5114. Claude Opus 5 high$0.5515. GPT-6 Astra xhigh$0.5516. Claude Fable 5.1 high$0.7317. Claude Opus 5 xhigh$0.7318. GPT-6 Astra max$0.7619. Claude Opus 5 max$0.9420. Claude Fable 5.1 xhigh$1.3221. Claude Fable 5.1 max$2.04the 21 qualifying configurations by cost per solvehighlighted: labs outside Anthropic, OpenAI and Google
Open-weight rows are marked on the ranking; their price is the creator's own rate, which third-party hosts routinely undercut

When to leave it

Three triggers. Everything else stays on the everyday model.

1The reasoning is genuinely hard.
CritPt, research-level physics, still separates tiers sharply. Two cheap configurations elsewhere on the board score exactly zero on it while still billing for the attempt. Passing GPQA Diamond does not predict passing CritPt, so a high headline score is not a licence to send it your hardest work.
2The model must state facts it cannot look up.
That is what the honesty column measures. Plus 31 means that across 100 questions the model finished 31 correct answers ahead of its confident mistakes. Minus 10.3 means it finished about 10 confident mistakes behind. Every GPT-5.6 Luna configuration sits below zero, so none qualifies at any effort. Put retrieval in front of the model and the gate relaxes, because it is no longer asserting from memory.
3The task will run long.
Fewer than one prompt in ten runs past 15 minutes, but those prompts carry 29% of coding spend and 50% of knowledge-work spend. This index scores bounded work only and does not adjudicate what to escalate to. Route on predicted duration, not on prompt count.

How we measured

Public benchmarks, one costing method throughout, and a rule stated in advance rather than fitted to the answer.

Abstract

We rank 71 model configurations from 30 models and 11 labs on the cost of one correct answer, using seven public component evaluations of the Artificial Analysis Intelligence Index and first-party prices, retrieved 1 September 2026, with GPT-6 Astra from pages retrieved 4 September 2026. A configuration enters the ranking only if it clears a knowledge-honesty gate by more than measurement noise and holds at least 85% of the best score across six accuracy evaluations; 21 do. Cost per solve is cost per attempt divided by share solved, combined by geometric mean so no single benchmark dominates. We publish the sensitivity of the qualifying set to the bar and to the benchmark set, and every figure traces to a dated snapshot and versioned scripts.

1

Scope and inclusion

Every model on the Artificial Analysis leaderboard carrying a measured Intelligence Index score and a score on all seven component evaluations used here. A configuration missing even one benchmark cannot be scored the way the others are. Imputing the gap would mean inventing a number, and averaging over what happens to be present would quietly reward models for the benchmarks they skipped.

The rule is mechanical on purpose. Anyone can run it against the leaderboard and get the same list, which is what makes 'what is missing' an answerable question rather than a matter of who we happened to think of.

  • In practice it is a Terminal-Bench rule. In practice the coverage rule is a Terminal-Bench v2.1 rule. It is the newest of the seven and Artificial Analysis has run it on 229 of 637 records, so it is the single missing evaluation for 395 of the 434 models that cannot be scored here. Whole labs are excluded by it alone, Amazon's Nova line and Microsoft's Phi among them. That is a property of benchmark coverage, not of those models.
  • This is not the Intelligence Index. The Artificial Analysis Intelligence Index weights nine evaluations. The two omitted here, GDPval-AA at 20% and tau3-Banking at 14%, are long-horizon agentic work and out of scope by design. They are also the expensive ones: GDPval alone can be around 59% of a model's published cost per index task. Cost per solve here is therefore a different and much smaller number than the cost per task Artificial Analysis publishes, and the two should never be quoted against each other.
  • Muse Spark 1.3 at max effort. Scored on all seven evaluations and would rank near the top, but Artificial Analysis publishes no price for it at all and lists no serving host. A configuration with no price cannot enter a cost ranking. Its xhigh sibling is priced and does appear.
  • Four non-reasoning variants of Kimi and DeepSeek. Missing Terminal-Bench v2.1 entirely, and their index is flagged estimated rather than measured.
  • Five vendor-deprecated configurations. Superseded by a newer release. This is a buying guide, so a model you cannot adopt going forward is out, though it is named here rather than quietly dropped.
2

Benchmarks

Seven component evaluations of the Artificial Analysis Intelligence Index v4.1.1. Six are scored for accuracy and enter cost per solve; the seventh, AA-Omniscience, is the honesty gate and is never averaged in.

Table 1. Benchmarks in scope, what each measures, and how much it separates the field.

BenchmarkWhat it measures, and why it is in scope
GPQA DiamondGraduate-level science, multiple choice. The floor check. Nearly every current model clears it, so it separates almost nothing.saturated
HLEHumanity's Last Exam: hard closed-ended questions across many fields. Still separates tiers sharply. One of the three that carry real signal.discriminates
CritPtResearch-level physics reasoning. The hardest test in scope. Two cheap configurations score exactly zero while still billing.discriminates
SciCodeScientific coding, graded on sub-problems. Compressed. Ranks three through forty span nine points.saturated
Terminal-Bench v2.1Short agentic coding in a terminal, minutes per task. The one borderline inclusion: agentic, but bounded. Removing it shifts costs about 15% and does not change the ranking.discriminates
AA-LCRReasoning over documents around 100,000 tokens. Compressed. The cheap tier is genuinely competitive here.saturated
AA-OmniscienceKnowledge, with a penalty for confident errors. Not scored for accuracy. It is the honesty gate.the gate
3

Metrics

3.1

Cost per solve

Take the price of one attempt on a benchmark and divide it by the share of problems the configuration got right. That is what a correct answer costs when you can spot a miss and retry. Do this on each of the six accuracy benchmarks, then combine them with a geometric mean so no single hard or expensive benchmark dominates.

cost per solvee = cost per attempte ÷ share solvede
index cost per solve = geometric mean over the six accuracy evaluations e
Worked example
GPT-5.6 Sol at medium solves 92.6% of GPQA Diamond and an attempt costs $0.0157, so one correct answer costs $0.0157 divided by 0.926, which is $0.017. Repeat across the other five benchmarks and take the geometric mean: $0.14 per solve. Claude Sonnet 5 at max costs less per token but solves fewer problems per attempt, so it lands at $1.33.
3.2

Efficacy

On each of the six accuracy benchmarks, divide this configuration's score by the highest score any configuration reached on that benchmark. Average those six shares. 100% means it matched the leader everywhere.

3.3

Honesty and the noise band

On the AA-Omniscience knowledge test a correct answer scores plus one, a confident wrong answer minus one, and declining to answer zero. The total is expressed per 100 questions, so the scale runs from minus 100 to plus 100. Passing means above zero by more than the noise band described next.

AA-Omniscience is 6,000 questions, each scoring plus one, minus one or zero. The net is a mean of bounded scores, so its standard error is at most about 1.3 points on the published per-100 scale. A configuration within two standard errors of zero is reported as too close to call rather than passed or failed, because the test cannot separate it from chance. Six configurations sit in that band, and one of them, GPT-5.6 Terra at max effort, nets plus 0.05, which is four hundredths of a standard error from zero.

3.4

The 85% bar

A stated target, not a fitted one. An elbow-detection rule was tried first and rejected: with four to eight frontier points it is decided by whichever two neighbours happen to tie.

4

Results

Of 71 configurations, 21 clear both gates and are ranked. GPT-5.6 Sol at medium ranks first at $0.14 per solve with 86% efficacy, followed by GPT-6 Astra at low at $0.17 and GPT-5.6 Sol at high at $0.19. The qualifying set spans 5 labs (Anthropic, Google, Meta, Moonshot and OpenAI) and a 15-fold cost range, from $0.14 to $2.04, for 12 points of efficacy (Figure 1).

Of the rest, 20 pass the honesty gate but fall under the 85% bar, 6 sit inside the noise band on honesty and are not ranked, 20 are confidently wrong more often than right, and 4 score zero on at least one benchmark, which leaves cost per solve undefined. The cheapest configuration measured, GPT-5.6 Luna at low at $0.013 per solve, is in the failing tier; without the gate it would rank first.

Figure 1. The 21 qualifying configurations by cost per solve; labs outside Anthropic, OpenAI and Google highlighted.

1. GPT-5.6 Sol medium$0.142. GPT-6 Astra low$0.173. GPT-5.6 Sol high$0.194. Muse Spark 1.3 xhigh$0.215. Gemini 3.8 Flash high$0.226. GPT-6 Astra medium$0.267. GPT-5.6 Sol xhigh$0.268. GPT-6 Astra high$0.369. Claude Opus 5 medium$0.3810. Claude Fable 5.1 low$0.4011. GPT-5.6 Sol max$0.4112. Kimi K3 max$0.5013. Claude Fable 5.1 medium$0.5114. Claude Opus 5 high$0.5515. GPT-6 Astra xhigh$0.5516. Claude Fable 5.1 high$0.7317. Claude Opus 5 xhigh$0.7318. GPT-6 Astra max$0.7619. Claude Opus 5 max$0.9420. Claude Fable 5.1 xhigh$1.3221. Claude Fable 5.1 max$2.04
5

Sensitivity

Two judgement calls set who appears: the 85% bar and the benchmark set. Both move the answer, so both are published rather than footnoted.

Table 2. Configurations and labs that qualify as the efficacy bar moves; the honesty gate is held fixed.

Efficacy barQualifyLabs represented
90.0%14Anthropic, Meta, OpenAI
87.5%19Anthropic, Meta, Moonshot, OpenAI
85.0% (used here)21Anthropic, Google, Meta, Moonshot, OpenAI
82.5%27Anthropic, Google, Meta, Moonshot, OpenAI, Z.ai, xAI
80.0%33Alibaba, Anthropic, Google, Meta, Moonshot, OpenAI, Z.ai, xAI
75.0%37Alibaba, Anthropic, Google, Meta, Moonshot, OpenAI, Z.ai, xAI

One benchmark carries most of that. Remove CritPt and 13 more configurations clear the same bar, from 8 labs instead of 5. Research-level physics is where these models separate, and whether it belongs in your definition of everyday work is a judgement you should make rather than inherit.

6

Limitations

  • Data collected from benchmarks as of 1 September 2026. GPT-6 Astra was released 3 September 2026 and added from pages retrieved 4 September 2026. There are new releases and re-grades between index patches, so figures move. Zaun Research will update this index regularly.
  • The retry assumption. Cost per solve assumes a failed attempt can be detected and retried. That holds on graded, bounded work, which is the scope here. It does not hold on long agentic runs, which is one reason those are out of scope.
  • Per-token price is not in the ranking. It does not predict cost per solve. One model here lists at 40% of another per token and costs more per task, because it uses more tokens and more turns.
  • Effort settings are not comparable across models. One vendor’s medium is not another’s. Read a configuration as a whole, never the effort label on its own. GPT-5.6 Sol at medium and Claude Opus 5 at medium list within 20% of each other per output token ($20 against $25 per million), yet one CritPt attempt costs $0.13 on the first and $1.40 on the second. The gap is tokens spent per task, not price per token.
  • Open-weight rows are priced at the creator’s rate, which is the expensive end. Marked open. Third-party hosts frequently undercut it, in one case by three to four times, so those configurations are likely cheaper in practice than shown here. The two NVIDIA rows are different again: NVIDIA offers no first-party API, so they carry a cross-provider median and are marked median price.
  • Platform premiums are not applied. Prices are first-party. Buying the same model through a cloud marketplace can add materially to it.
  • Long-running work is out of scope. Every benchmark here is bounded. The escalation triggers above point at the tail; this index does not rank models for it.
7

Data and reproducibility

Every figure on this page is computed from a dated snapshot of public sources, and the scripts that turn the snapshot into the ranking are versioned with it. Per-evaluation scores and cost per attempt were read from the Artificial Analysis evaluation pages [3–9] and checked against the published cost per task on each; index composition and weights from the Artificial Analysis methodology [2]; model-level cost per index task and per-token prices from the Artificial Analysis model pages [1]. First-party per-token prices for Anthropic and OpenAI were confirmed against the vendors’ own pricing pages [10, 11]; the cloud-platform multipliers discussed in the limitations come from the Bedrock, Azure and Vertex price lists [12–14]. Every other vendor’s first-party price is as recorded by Artificial Analysis on that model’s page [1]. The task-length figures are Zaun internal telemetry [15], August 2026, one prompt measured from submission to its last completion event, published in aggregate only. This is version 0.1 of the index, a preliminary snapshot; later versions will cite the snapshot date they supersede.

References

  1. Artificial Analysis, model pages and Intelligence Index leaderboard, v4.1.1. artificialanalysis.ai/models. Retrieved 1 September 2026; GPT-6 Astra pages 4 September 2026.
  2. Artificial Analysis, Intelligence benchmarking methodology. artificialanalysis.ai/methodology/intelligence-benchmarking.
  3. Rein, D. et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark, 2023, arXiv:2311.12022. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/gpqa-diamond.
  4. Phan, L. et al., Humanity’s Last Exam, 2025, arXiv:2501.14249. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/humanitys-last-exam.
  5. Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark, 2025, arXiv:2509.26574. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/critpt.
  6. Tian, M. et al., SciCode: A Research Coding Benchmark Curated by Scientists, 2024, arXiv:2407.13168. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/scicode.
  7. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces, 2026, arXiv:2601.11868; tbench.ai, artificialanalysis.ai/evaluations/terminalbench-v2-1. Version 2.1 scored by Artificial Analysis.
  8. Artificial Analysis, Long Context Reasoning (AA-LCR). artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning.
  9. Artificial Analysis, AA-Omniscience: knowledge with a penalty for confident errors. artificialanalysis.ai/evaluations/omniscience.
  10. Anthropic, Claude pricing. platform.claude.com/docs/en/about-claude/pricing. Retrieved 1 September 2026.
  11. OpenAI, API pricing. developers.openai.com/api/docs/pricing. Retrieved 1 September 2026.
  12. Amazon Web Services, Amazon Bedrock pricing. aws.amazon.com/bedrock/pricing. Retrieved 1 September 2026.
  13. Microsoft, Azure Retail Prices API, Foundry meters, eastus2. prices.azure.com/api/retail/prices. Retrieved 1 September 2026.
  14. Google Cloud, Vertex AI generative AI pricing. cloud.google.com/vertex-ai/generative-ai/pricing. Retrieved 1 September 2026.
  15. Zaun Research, task-length distribution per prompt, coding agents and knowledge work, August 2026. Internal telemetry, aggregate figures only.