Skip to content

Choosing the everyday model.

A method for choosing an agent’s default model, using knowledge honesty, benchmark efficacy, and cost per correct answer.

The research behind the Everyday Model Index.

Most work is short

Across our own agent telemetry, more than nine in ten prompts resolve inside fifteen minutes. These are the tasks that make up much of the observed workload. Duration motivates evaluating a less expensive default; it does not, by itself, establish which model can complete a task.

Short by count. Uneven by cost.

Green marks work completed in under 15 minutes. Violet marks the longer-running tail.

Coding agents

17,609 measured prompts
Prompt count
94.4%under 15 minutes
Token spend
71.2%under 15 minutes

5.6% of prompts account for 28.8% of spend.

Knowledge work

3,796 measured prompts
Prompt count
91.9%under 15 minutes
Token spend
49.7%under 15 minutes

8.1% of prompts account for 50.3% of spend.

Each square represents approximately one percentage point. Aggregate observations motivate examining a lower-cost default; they do not establish savings from rerouting these prompts.
Where the saving is, and where it is not

Those sub-15-minute prompts carry 71% of coding spend and 50% of knowledge-work spend. At the two ends of the qualifying set, a correct answer costs $0.14 on the cheapest qualifier against $2.04 on Claude Fable 5.1 at max, a 15× difference at 86% efficacy against 98%. These are benchmark costs, not measured savings from rerouting those prompts.

The caveat is the tail. It is small by count and large by cost, larger in knowledge work than in coding, so an everyday model is not a licence to stop paying attention to it. The escalation rule below handles it. Route on predicted duration, not on prompt count, or the saving will disappoint.

Zaun internal data · August 2026 · a prompt is measured from submission to its last completion event

Three questions

Every configuration is asked the same three, in this order. Fail the first and cost never matters; fail the second and it does not reach the ranking.

Qualification comes before price.

GPT-5.6 Sol (medium) clears both gates and enters the cost ranking.

Knowledge honesty+19.4Net correct per 100 questions

Above the +2.58 noise-band boundary.

Benchmark efficacy85.5%Combined share of the best scores

Above the published 85% threshold.

Cost ranking$0.14Per correct answer

Compare with the other qualifying configurations.

Fails honestyGPT-5.6 Luna (low)

−14.7 net correct per 100 questions. Low cost cannot override failed honesty.

Below the efficacy barGPT-5.6 Sol (low)

77.8% efficacy. Honesty passes, but capability falls below the bar.

Honesty inconclusiveGPT-5.6 Sol (non-reasoning)

+1.1 net correct per 100 questions. The margin stays inside the noise band.

Recorded configurations from the index snapshot. The diagram explains a selection rule, not a running router.
How the three measures are defined
First

Is it honest?

Right answers minus confidently wrong ones. Below zero it asserts more than it knows. Confident errors are paid for in labor hours, not tokens, so a model that asserts more than it knows is out before cost is even considered.

Second

Is it capable enough?

How close it gets to the best score on each test, averaged across six benchmarks. The six are graduate and research-level science, frontier exams, scientific coding, terminal work and long-document reasoning. We use 85% of the best score across them as the bar to enter the ranking. This is a benchmark threshold, not a guarantee for every workload.

Third

What does it cost?

What one correct answer costs once you have paid for the attempts that failed. This is what the ranking sorts on.

What the data shows

GPT-5.6 Sol

Reasoning effort changes the configuration. The lowest settings here fail a gate before they can enter the ranking.

Cheapest qualifying effort: medium$0.1485.5% efficacy
Maximum effort$0.4193.6% efficacy
Cost and efficacy across GPT-5.6 Sol effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted medium: $0.14, 85.5% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.25$0.5$1$2Cost per correct answer (USD, log scale)Cost and efficacy across GPT-5.6 Sol effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted medium: $0.14, 85.5% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.5$2Cost / solve (USD, log scale)
3.0× the cost per correct answer at max effort.

Efficacy changes by 8.1 percentage points. The filled point marks the cheapest qualifying setting; the dashed line marks the 85% bar.

Recorded configurations connected in effort order. Cost uses a log scale. The plot does not measure latency or intermediate settings.

Claude Opus 5

The same comparison within another model family. Effort labels are specific to each vendor.

Cheapest qualifying effort: medium$0.3889.3% efficacy
Maximum effort$0.9493.0% efficacy
Cost and efficacy across Claude Opus 5 effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted medium: $0.38, 89.3% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.25$0.5$1$2Cost per correct answer (USD, log scale)Cost and efficacy across Claude Opus 5 effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted medium: $0.38, 89.3% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.5$2Cost / solve (USD, log scale)
2.5× the cost per correct answer at max effort.

Efficacy changes by 3.7 percentage points. The filled point marks the cheapest qualifying setting; the dashed line marks the 85% bar.

Recorded configurations connected in effort order. Cost uses a log scale. The plot does not measure latency or intermediate settings.

Claude Fable 5.1

The same comparison within another model family. Effort labels are specific to each vendor.

Cheapest qualifying effort: low$0.4089.7% efficacy
Maximum effort$2.0497.5% efficacy
Cost and efficacy across Claude Fable 5.1 effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted low: $0.40, 89.7% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.25$0.5$1$2Cost per correct answer (USD, log scale)Cost and efficacy across Claude Fable 5.1 effort settingsHorizontal axis: cost per correct answer in dollars on a log scale. Vertical axis: benchmark efficacy. Highlighted low: $0.40, 89.7% efficacy. The dashed line is the 85% threshold. The full configuration data is available in the Everyday Model Index.50%70%100%85%Efficacy$0.1$0.5$2Cost / solve (USD, log scale)
5.1× the cost per correct answer at max effort.

Efficacy changes by 7.9 percentage points. The filled point marks the cheapest qualifying setting; the dashed line marks the 85% bar.

Recorded configurations connected in effort order. Cost uses a log scale. The plot does not measure latency or intermediate settings.

Four findings from the same snapshot. Each includes the comparison and its supporting evidence.

3.3×Use the reasoning lever, but keep the big model.

Claude Sonnet 5 at max effort costs less per token than Claude Fable 5.1, and 3.3 times more per correct answer, because it uses more tokens and solves fewer problems per attempt. Fable 5.1 at low effort is cheaper and 9 points more capable. The sticker price told the opposite story.

Claude Sonnet 5 max$1.3381% efficacyClaude Opus 5 medium$0.3889% efficacyClaude Fable 5.1 low$0.4090% efficacycost per solve, same vendor, all pass the honesty gate
Sonnet 5 max $1.33 per solve at 81% efficacy · Fable 5.1 low $0.40 at 90% · same vendor, both pass the honesty gate
20 of 71The cheapest configurations are among the ones that make things up.

20 configurations are right less often than they are confidently wrong. GPT-5.6 Luna (low) would top the whole ranking at $0.013 per solve if the gate did not exist. It is excluded because every confident error is undone by a person, at a labor rate that swamps the saving. These models look fine on a standard test: their median GPQA Diamond score is 90%. Honesty is not something you can read off accuracy.

noise band−40−20+0+20+40$0.03$0.1$0.3$1every GPT-5.6 Luna settingGPT-5.6 Sol medium, cheapest to qualifycost per solve, log scale · vertical: honesty margin per 100 questions
Green passes, amber is within measurement noise of zero, red fails · every GPT-5.6 Luna setting fails, best margin −10.3
The effort setting moves cost more than the model name does.

GPT-5.6 Sol alone spans $0.14 to $0.41 per solve and 86% to 94% efficacy across its effort settings. That is most of the useful range without changing model. The same is true of every frontier family here, which is why a configuration, model plus effort, is the unit ranked and never the model alone.

Sol at low (78%) falls under the bar; at non-reasoning its honesty is too close to call · the setting matters on both gates
1 of 10Labs outside the big three hold top-ten places.

Muse Spark 1.3 ranks 4. Of the 21 configurations that qualify, one has open weights and can be self-hosted. The field is wider than the three names most buying conversations start with, and it is moving fast: GPT-6 Astra, released 3 September 2026, was added on 4 September 2026 and holds 3 of the top 10 places.

1. GPT-5.6 Sol medium$0.142. GPT-6 Astra low$0.173. GPT-5.6 Sol high$0.194. Muse Spark 1.3 xhigh$0.215. Gemini 3.8 Flash high$0.226. GPT-6 Astra medium$0.267. GPT-5.6 Sol xhigh$0.268. GPT-6 Astra high$0.369. Claude Opus 5 medium$0.3810. Claude Fable 5.1 low$0.4011. GPT-5.6 Sol max$0.4112. Kimi K3 max$0.5013. Claude Fable 5.1 medium$0.5114. Claude Opus 5 high$0.5515. GPT-6 Astra xhigh$0.5516. Claude Fable 5.1 high$0.7317. Claude Opus 5 xhigh$0.7318. GPT-6 Astra max$0.7619. Claude Opus 5 max$0.9420. Claude Fable 5.1 xhigh$1.3221. Claude Fable 5.1 max$2.04the 21 qualifying configurations by cost per solvehighlighted: labs outside Anthropic, OpenAI and Google
Open-weight rows are marked in the index; their price is the creator's own rate, which third-party hosts routinely undercut

When to leave it

Three conditions to evaluate when deciding whether to leave the default route.

1The reasoning is hard.
CritPt, research-level physics, still separates tiers sharply. Two cheap configurations elsewhere on the board score exactly zero on it while still billing for the attempt. Passing GPQA Diamond does not predict passing CritPt, so a high headline score is not a licence to send it your hardest work.
2The model must state facts it cannot look up.
That is what the honesty column measures. Plus 31 means that across 100 questions the model finished 31 correct answers ahead of its confident mistakes. Minus 10.3 means it finished about 10 confident mistakes behind. Every GPT-5.6 Luna configuration sits below zero, so none qualifies at any effort. Put retrieval in front of the model and the gate relaxes, because it is no longer asserting from memory.
3The task will run long.
Fewer than one prompt in ten runs past 15 minutes, but those prompts carry 29% of coding spend and 50% of knowledge-work spend. The index scores bounded work only and does not adjudicate what to escalate to. Route on predicted duration, not on prompt count.

How we measured

Public benchmarks, one costing method throughout, and a rule stated in advance rather than fitted to the answer.

Abstract

We rank 71 model configurations from 30 models and 11 labs on the cost of one correct answer, using seven public component evaluations of the Artificial Analysis Intelligence Index and first-party prices, retrieved 1 September 2026, with GPT-6 Astra from pages retrieved 4 September 2026. A configuration enters the ranking only if it clears a knowledge-honesty gate by more than measurement noise and holds at least 85% of the best score across six accuracy evaluations; 21 do. Cost per solve is cost per attempt divided by share solved, combined by geometric mean so no single benchmark dominates. We publish the sensitivity of the qualifying set to the bar and to the benchmark set, and every figure traces to a dated snapshot and versioned scripts.

1

Scope and inclusion

Every model on the Artificial Analysis leaderboard carrying a measured Intelligence Index score and a score on all seven component evaluations used here. A configuration missing even one benchmark cannot be scored the way the others are. Imputing the gap would mean inventing a number, and averaging over what happens to be present would quietly reward models for the benchmarks they skipped.

The rule is mechanical on purpose. Anyone can run it against the leaderboard and get the same list, which is what makes 'what is missing' an answerable question rather than a matter of who we happened to think of.

  • In practice it is a Terminal-Bench rule. In practice the coverage rule is a Terminal-Bench v2.1 rule. It is the newest of the seven and Artificial Analysis has run it on 229 of 637 records, so it is the single missing evaluation for 395 of the 434 models that cannot be scored here. Whole labs are excluded by it alone, Amazon's Nova line and Microsoft's Phi among them. That is a property of benchmark coverage, not of those models.
  • This is not the Intelligence Index. The Artificial Analysis Intelligence Index weights nine evaluations. The two omitted here, GDPval-AA at 20% and tau3-Banking at 14%, are long-horizon agentic work and out of scope by design. They are also the expensive ones: GDPval alone can be around 59% of a model's published cost per index task. Cost per solve here is therefore a different and much smaller number than the cost per task Artificial Analysis publishes, and the two should never be quoted against each other.
  • Muse Spark 1.3 at max effort. Scored on all seven evaluations and would rank near the top, but Artificial Analysis publishes no price for it at all and lists no serving host. A configuration with no price cannot enter a cost ranking. Its xhigh sibling is priced and does appear.
  • Four non-reasoning variants of Kimi and DeepSeek. Missing Terminal-Bench v2.1 entirely, and their index is flagged estimated rather than measured.
  • Five vendor-deprecated configurations. Superseded by a newer release. This is a buying guide, so a model you cannot adopt going forward is out, though it is named here rather than quietly dropped.
2

Benchmarks

Seven component evaluations of the Artificial Analysis Intelligence Index v4.1.1. Six are scored for accuracy and enter cost per solve; the seventh, AA-Omniscience, is the honesty gate and is never averaged in.

Table 1. Benchmarks in scope, what each measures, and how much it separates the field.

BenchmarkWhat it measures, and why it is in scope
GPQA DiamondGraduate-level science, multiple choice. The floor check. Nearly every current model clears it, so it separates almost nothing.saturated
HLEHumanity's Last Exam: hard closed-ended questions across many fields. Still separates tiers sharply. One of the three that carry real signal.discriminates
CritPtResearch-level physics reasoning. The hardest test in scope. Two cheap configurations score exactly zero while still billing.discriminates
SciCodeScientific coding, graded on sub-problems. Compressed. Ranks three through forty span nine points.saturated
Terminal-Bench v2.1Short agentic coding in a terminal, minutes per task. The one borderline inclusion: agentic, but bounded. Removing it shifts costs about 15% and does not change the ranking.discriminates
AA-LCRReasoning over documents around 100,000 tokens. Compressed. The cheap tier is genuinely competitive here.saturated
AA-OmniscienceKnowledge, with a penalty for confident errors. Not scored for accuracy. It is the honesty gate.the gate
3

Metrics

3.1

Cost per solve

Take the price of one attempt on a benchmark and divide it by the share of problems the configuration got right. That is what a correct answer costs when you can spot a miss and retry. Do this on each of the six accuracy benchmarks, then combine them with a geometric mean so no single hard or expensive benchmark dominates.

cost per solvee = cost per attempte ÷ share solvede
index cost per solve = geometric mean over the six accuracy evaluations e
Worked example
GPT-5.6 Sol at medium solves 92.6% of GPQA Diamond and an attempt costs $0.0157, so one correct answer costs $0.0157 divided by 0.926, which is $0.017. Repeat across the other five benchmarks and take the geometric mean: $0.14 per solve. Claude Sonnet 5 at max costs less per token but solves fewer problems per attempt, so it lands at $1.33.
3.2

Efficacy

On each of the six accuracy benchmarks, divide this configuration's score by the highest score any configuration reached on that benchmark. Average those six shares. 100% means it matched the leader everywhere.

3.3

Honesty and the noise band

On the AA-Omniscience knowledge test a correct answer scores plus one, a confident wrong answer minus one, and declining to answer zero. The total is expressed per 100 questions, so the scale runs from minus 100 to plus 100. Passing means above zero by more than the noise band described next.

AA-Omniscience is 6,000 questions, each scoring plus one, minus one or zero. The net is a mean of bounded scores, so its standard error is at most about 1.3 points on the published per-100 scale. A configuration within two standard errors of zero is reported as too close to call rather than passed or failed, because the test cannot separate it from chance. Six configurations sit in that band, and one of them, GPT-5.6 Terra at max effort, nets plus 0.05, which is four hundredths of a standard error from zero.

3.4

The 85% bar

A stated target, not a fitted one. An elbow-detection rule was tried first and rejected: with four to eight frontier points it is decided by whichever two neighbours happen to tie.

4

Results

Of 71 configurations, 21 clear both gates and are ranked. GPT-5.6 Sol at medium ranks first at $0.14 per solve with 86% efficacy, followed by GPT-6 Astra at low at $0.17 and GPT-5.6 Sol at high at $0.19. The qualifying set spans 5 labs (Anthropic, Google, Meta, Moonshot and OpenAI) and a 15-fold cost range, from $0.14 to $2.04, for 12 points of efficacy (Figure 1).

Of the rest, 20 pass the honesty gate but fall under the 85% bar, 6 sit inside the noise band on honesty and are not ranked, 20 are confidently wrong more often than right, and 4 score zero on at least one benchmark, which leaves cost per solve undefined. The cheapest configuration measured, GPT-5.6 Luna at low at $0.013 per solve, is in the failing tier; without the gate it would rank first.

Figure 1. The 21 qualifying configurations by cost per solve; labs outside Anthropic, OpenAI and Google highlighted.

1. GPT-5.6 Sol medium$0.142. GPT-6 Astra low$0.173. GPT-5.6 Sol high$0.194. Muse Spark 1.3 xhigh$0.215. Gemini 3.8 Flash high$0.226. GPT-6 Astra medium$0.267. GPT-5.6 Sol xhigh$0.268. GPT-6 Astra high$0.369. Claude Opus 5 medium$0.3810. Claude Fable 5.1 low$0.4011. GPT-5.6 Sol max$0.4112. Kimi K3 max$0.5013. Claude Fable 5.1 medium$0.5114. Claude Opus 5 high$0.5515. GPT-6 Astra xhigh$0.5516. Claude Fable 5.1 high$0.7317. Claude Opus 5 xhigh$0.7318. GPT-6 Astra max$0.7619. Claude Opus 5 max$0.9420. Claude Fable 5.1 xhigh$1.3221. Claude Fable 5.1 max$2.04
5

Sensitivity

A stricter bar narrows the field.

The honesty gate and six benchmarks stay fixed. Only the efficacy threshold changes.

Qualifying configurations decline from 37 at 75% efficacy to 14 at 90%. The published 85% bar admits 21 configurations.02040Qualifying configurations37332721191475%85%90%Efficacy threshold
Published rule / 85%21

qualifying configurations
from 5 labs

GPT-5.6 Sol (medium)

Cheapest qualifier at $0.14 per correct answer.

The six points reproduce the published sensitivity sweep. The full threshold values and labs appear in Table 2 below.

Two judgement calls set who appears: the 85% bar and the benchmark set. Both move the answer, so both are published rather than footnoted.

Table 2. Configurations and labs that qualify as the efficacy bar moves; the honesty gate is held fixed.

Efficacy barQualifyLabs represented
90.0%14Anthropic, Meta, OpenAI
87.5%19Anthropic, Meta, Moonshot, OpenAI
85.0% (used here)21Anthropic, Google, Meta, Moonshot, OpenAI
82.5%27Anthropic, Google, Meta, Moonshot, OpenAI, Z.ai, xAI
80.0%33Alibaba, Anthropic, Google, Meta, Moonshot, OpenAI, Z.ai, xAI
75.0%37Alibaba, Anthropic, Google, Meta, Moonshot, OpenAI, Z.ai, xAI

One benchmark carries most of that. Remove CritPt and 13 more configurations clear the same bar, from 8 labs instead of 5. Research-level physics is where these models separate, and whether it belongs in your definition of everyday work is a judgement you should make rather than inherit.

6

Limitations

The index informs model selection; it does not evaluate an end-to-end router. Latency, tool access, context limits, retry behavior, and escalation policy still need to be tested on the workload being routed. Cost per solve is a benchmark estimate, not a measured production bill.

  • Data collected from benchmarks as of 1 September 2026. GPT-6 Astra was released 3 September 2026 and added from pages retrieved 4 September 2026. There are new releases and re-grades between index patches, so figures move. Zaun Research will update this index regularly.
  • The retry assumption. Cost per solve assumes a failed attempt can be detected and retried. That holds on graded, bounded work, which is the scope here. It does not hold on long agentic runs, which is one reason those are out of scope.
  • Per-token price is not in the ranking. It does not predict cost per solve. One model here lists at 40% of another per token and costs more per task, because it uses more tokens and more turns.
  • Effort settings are not comparable across models. One vendor’s medium is not another’s. Read a configuration as a whole, never the effort label on its own. GPT-5.6 Sol at medium and Claude Opus 5 at medium list within 20% of each other per output token ($20 against $25 per million), yet one CritPt attempt costs $0.13 on the first and $1.40 on the second. The gap is tokens spent per task, not price per token.
  • Open-weight rows are priced at the creator’s rate, which is the expensive end. Marked open. Third-party hosts frequently undercut it, in one case by three to four times, so those configurations are likely cheaper in practice than shown here. The two NVIDIA rows are different again: NVIDIA offers no first-party API, so they carry a cross-provider median and are marked median price.
  • Platform premiums are not applied. Prices are first-party. Buying the same model through a cloud marketplace can add materially to it.
  • Long-running work is out of scope. Every benchmark here is bounded. The escalation triggers above point at the tail; this index does not rank models for it.
7

Data and reproducibility

Every figure on this page is computed from a dated snapshot of public sources, and the scripts that turn the snapshot into the ranking are versioned with it. Per-evaluation scores and cost per attempt were read from the Artificial Analysis evaluation pages [3-9] and checked against the published cost per task on each; index composition and weights from the Artificial Analysis methodology [2]; model-level cost per index task and per-token prices from the Artificial Analysis model pages [1]. First-party per-token prices for Anthropic and OpenAI were confirmed against the vendors’ own pricing pages [10, 11]; the cloud-platform multipliers discussed in the limitations come from the Bedrock, Azure and Vertex price lists [12-14]. Every other vendor’s first-party price is as recorded by Artificial Analysis on that model’s page [1]. The task-length figures are Zaun internal telemetry [15], August 2026, one prompt measured from submission to its last completion event, published in aggregate only. This is version 0.1 of the index, a preliminary snapshot; later versions will cite the snapshot date they supersede.

References

  1. Artificial Analysis, model pages and Intelligence Index leaderboard, v4.1.1. artificialanalysis.ai/models. Retrieved 1 September 2026; GPT-6 Astra pages 4 September 2026.
  2. Artificial Analysis, Intelligence benchmarking methodology. artificialanalysis.ai/methodology/intelligence-benchmarking.
  3. Rein, D. et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark, 2023, arXiv:2311.12022. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/gpqa-diamond.
  4. Phan, L. et al., Humanity’s Last Exam, 2025, arXiv:2501.14249. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/humanitys-last-exam.
  5. Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark, 2025, arXiv:2509.26574. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/critpt.
  6. Tian, M. et al., SciCode: A Research Coding Benchmark Curated by Scientists, 2024, arXiv:2407.13168. Scored by Artificial Analysis, artificialanalysis.ai/evaluations/scicode.
  7. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces, 2026, arXiv:2601.11868; tbench.ai, artificialanalysis.ai/evaluations/terminalbench-v2-1. Version 2.1 scored by Artificial Analysis.
  8. Artificial Analysis, Long Context Reasoning (AA-LCR). artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning.
  9. Artificial Analysis, AA-Omniscience: knowledge with a penalty for confident errors. artificialanalysis.ai/evaluations/omniscience.
  10. Anthropic, Claude pricing. platform.claude.com/docs/en/about-claude/pricing. Retrieved 1 September 2026.
  11. OpenAI, API pricing. developers.openai.com/api/docs/pricing. Retrieved 1 September 2026.
  12. Amazon Web Services, Amazon Bedrock pricing. aws.amazon.com/bedrock/pricing. Retrieved 1 September 2026.
  13. Microsoft, Azure Retail Prices API, Foundry meters, eastus2. prices.azure.com/api/retail/prices. Retrieved 1 September 2026.
  14. Google Cloud, Vertex AI generative AI pricing. cloud.google.com/vertex-ai/generative-ai/pricing. Retrieved 1 September 2026.
  15. Zaun Research, task-length distribution per prompt, coding agents and knowledge work, August 2026. Internal telemetry, aggregate figures only.