Back to information library
ARTIFICIAL INTELLIGENCE / PUBLIC DATA

How do AI thinking levels affect cost, speed, and accuracy?

Compare reasoning levels across OpenAI GPT, Google Gemini, and Anthropic Claude, showing intelligence scores, task costs, and execution speeds.

Collected UTC
Coverage: 34 model reasoning tiers across OpenAI, Google, and Anthropic.

Observations

Best picks and recommendations

MOST EFFICIENT

Gemini 3.8 / 3.7 Flash Medium

Score 43-44 · $0.0020-$0.0022/task · Value Index: 20,000+

Top-tier math and coding at fraction-of-a-cent rates with under 10-second latency.

BEST BUDGET WORKHORSE

GPT-5.6 Luna High

Score 33 · $0.09/task · Value Index: 183

High-accuracy multi-step research and structured code generation.

BEST DEVELOPER BALANCE

Claude 3.7 Sonnet Low / Medium

Score 43-49 · $0.03-$0.07/task · Value Index: 700+

The premier sweet spot for daily software engineering, refactoring, and agent loops.

MAXIMUM INTELLIGENCE

Claude 3.7 Sonnet Max or GPT-6 Astra Max

Score 53-55 · $0.95-$3.26/task · Value Index: 16-58

Reserved strictly for mission-critical architecture, novel mathematics, and high-stakes proofs.

Visual comparisons: Intelligence versus Cost and Speed

Filter charts by provider:

Intelligence Index vs. Cost per Task

Higher and to the left represents superior efficiency (greater intelligence at lower task cost). Key anchor models are labeled directly on the chart dots.

View cost chart data table
ModelProviderLevelAvg Cost / TaskIntelligence Index
Claude 3.7 SonnetAnthropicHigh$0.220053
Claude 3.7 SonnetAnthropicLow$0.030043
Claude 3.7 SonnetAnthropicMax$0.950055
Claude 3.7 SonnetAnthropicMedium$0.070049
Claude 3.7 SonnetAnthropicStandard$0.010036
Claude Mythos 5.1AnthropicHigh$1.200054
Claude Mythos 5.1AnthropicMax$4.800056
Gemini 3.7 FlashGoogleMedium$0.002043
Gemini 3.7 ProGoogleHigh$0.050051
Gemini 3.7 ProGoogleMedium$0.025049
Gemini 3.8 FlashGoogleHigh$0.007548
Gemini 3.8 FlashGoogleLow$0.000935
Gemini 3.8 FlashGoogleMax$0.012050
Gemini 3.8 FlashGoogleMedium$0.002244
GPT-5.6 LunaOpenAIExtra High$0.120035
GPT-5.6 LunaOpenAIHigh$0.090033
GPT-5.6 LunaOpenAILow$0.040022
GPT-5.6 LunaOpenAIMax$0.180038
GPT-5.6 LunaOpenAIMedium$0.060026
GPT-5.6 SolOpenAIExtra High$1.430044
GPT-5.6 SolOpenAIHigh$0.810042
GPT-5.6 SolOpenAILow$0.230034
GPT-5.6 SolOpenAIMax$1.990047
GPT-5.6 SolOpenAIMedium$0.500039
GPT-5.6 TerraOpenAIExtra High$0.930038
GPT-5.6 TerraOpenAIHigh$0.520034
GPT-5.6 TerraOpenAILow$0.120028
GPT-5.6 TerraOpenAIMax$1.400042
GPT-5.6 TerraOpenAIMedium$0.330033
GPT-6 AstraOpenAIExtra High$3.260053
GPT-6 AstraOpenAIHigh$2.710051
GPT-6 AstraOpenAILow$0.820046
GPT-6 AstraOpenAIMax$3.260053
GPT-6 AstraOpenAIMedium$2.170050

Intelligence Index vs. Execution Speed & Latency

Comparing capability against response turnaround tiers: Fast (1-5s), Moderate (6-20s), and Deliberate (30-150s). Notice the sharp latency penalty required for top-tier reasoning effort.

View speed chart data table
ModelProviderLevelSpeed TierIntelligence Index
Claude 3.7 SonnetAnthropicHighDeliberate (30-150s)53
Claude 3.7 SonnetAnthropicLowModerate (6-20s)43
Claude 3.7 SonnetAnthropicMaxDeliberate (30-150s)55
Claude 3.7 SonnetAnthropicMediumModerate (6-20s)49
Claude 3.7 SonnetAnthropicStandardFast (1-5s)36
Claude Mythos 5.1AnthropicHighDeliberate (30-150s)54
Claude Mythos 5.1AnthropicMaxDeliberate (30-150s)56
Gemini 3.7 FlashGoogleMediumModerate (6-20s)43
Gemini 3.7 ProGoogleHighDeliberate (30-150s)51
Gemini 3.7 ProGoogleMediumModerate (6-20s)49
Gemini 3.8 FlashGoogleHighDeliberate (30-150s)48
Gemini 3.8 FlashGoogleLowFast (1-5s)35
Gemini 3.8 FlashGoogleMaxDeliberate (30-150s)50
Gemini 3.8 FlashGoogleMediumModerate (6-20s)44
GPT-5.6 LunaOpenAIExtra HighDeliberate (30-150s)35
GPT-5.6 LunaOpenAIHighDeliberate (30-150s)33
GPT-5.6 LunaOpenAILowFast (1-5s)22
GPT-5.6 LunaOpenAIMaxDeliberate (30-150s)38
GPT-5.6 LunaOpenAIMediumModerate (6-20s)26
GPT-5.6 SolOpenAIExtra HighDeliberate (30-150s)44
GPT-5.6 SolOpenAIHighDeliberate (30-150s)42
GPT-5.6 SolOpenAILowModerate (6-20s)34
GPT-5.6 SolOpenAIMaxDeliberate (30-150s)47
GPT-5.6 SolOpenAIMediumModerate (6-20s)39
GPT-5.6 TerraOpenAIExtra HighDeliberate (30-150s)38
GPT-5.6 TerraOpenAIHighDeliberate (30-150s)34
GPT-5.6 TerraOpenAILowModerate (6-20s)28
GPT-5.6 TerraOpenAIMaxDeliberate (30-150s)42
GPT-5.6 TerraOpenAIMediumModerate (6-20s)33
GPT-6 AstraOpenAIExtra HighDeliberate (30-150s)53
GPT-6 AstraOpenAIHighDeliberate (30-150s)51
GPT-6 AstraOpenAILowModerate (6-20s)46
GPT-6 AstraOpenAIMaxDeliberate (30-150s)53
GPT-6 AstraOpenAIMediumModerate (6-20s)50

About the data

This dataset compares reasoning capabilities, pricing, task costs, and execution speeds across frontier AI models from OpenAI, Anthropic, and Google. It provides practitioners and engineers with the trade-offs needed to configure reasoning effort levels in production software.

Key Metrics and Definitions

  • Intelligence Index (0-100): A normalized benchmark index combining competitive STEM mathematics (AIME 2024, MATH-500), expert scientific reasoning (GPQA Diamond), and real-world software engineering (SWE-bench Verified).
  • Token Pricing ($ / 1M tokens): Published direct provider API token rates for input and output, measured in USD per million tokens. Internal reasoning and thinking tokens are billed at the output rate.
  • Average Task Cost ($ / task): The estimated financial cost to complete a representative multi-step analytical prompt with 1,000 input tokens and typical reasoning token generation for that effort tier.
  • Speed: A categorical rating (Fast, Moderate, Deliberate, Low) reflecting end-to-end task completion latency in seconds.
  • Depth: The relative depth of the internal reasoning graph, tracking token budgets from short verification checks to exhaustive search.
  • Value Index: Calculated as normalized intelligence divided by average task cost, highlighting models that deliver the highest intelligence per dollar spent.

Coverage Limits

Prices and performance figures reflect official provider documentation and benchmark cards as of September 2026. Latency figures represent average direct API performance and may vary based on provider server load, region, and batching.

Possible explanations

Test-time compute scaling and search tree dynamics

Proposed mechanism: Reasoning models use reinforcement learning to search through multiple potential solution paths, backtrack when hitting dead ends, and verify steps internally before generating a final response.

Evidence: Performance on formal verification tasks (like competition mathematics and code execution) jumps significantly with thinking tokens. However, once the model finds the correct solution path, additional reasoning tokens provide diminishing returns, explaining the plateau between High and Max levels.

Asymmetric token pricing and output compute weight

Proposed mechanism: Output tokens require sequential autoregressive generation, which cannot be parallelized across tokens like prompt ingestion. Providers price output tokens 3x to 5x higher than input tokens to reflect hardware allocation.

Evidence: Because all thinking tokens are billed as output tokens, a query that generates 15,000 thinking tokens shifts over 90% of the total request cost onto the generation stage, causing exponential cost growth at high effort levels.

Explore the records

Search and compare reasoning models by provider, intelligence score, cost per task, and recommended use case.

Frontier model thinking levels comparison
ModelProviderReasoning LevelIntelligence IndexInput / 1MOutput / 1MAvg Cost / TaskSpeedDepthBest For
Claude 3.7 SonnetAnthropicHigh53$3.00$15.00$0.2200Deliberate (30-150s)Deep (3/4)Complex tasks and deep logic
Claude 3.7 SonnetAnthropicLow43$3.00$15.00$0.0300Moderate (6-20s)Check (1/4)High intelligence at balanced cost
Claude 3.7 SonnetAnthropicMax55$3.00$15.00$0.9500Deliberate (30-150s)Exhaustive (4/4)Hardest possible tasks
Claude 3.7 SonnetAnthropicMedium49$3.00$15.00$0.0700Moderate (6-20s)Standard (2/4)Complex everyday development
Claude 3.7 SonnetAnthropicStandard36$3.00$15.00$0.0100Fast (1-5s)Check (1/4)Everyday tasks with rapid responses
Claude Mythos 5.1AnthropicHigh54$15.00$75.00$1.2000Deliberate (30-150s)Deep (3/4)Frontier-level multi-step tasks
Claude Mythos 5.1AnthropicMax56$15.00$75.00$4.8000Deliberate (30-150s)Exhaustive (4/4)Maximum capability frontier science
Gemini 3.7 FlashGoogleMedium43$0.10$0.40$0.0020Moderate (6-20s)Standard (2/4)High-efficiency coding and problem solving
Gemini 3.7 ProGoogleHigh51$1.25$5.00$0.0500Deliberate (30-150s)Deep (3/4)Complex multimodal planning
Gemini 3.7 ProGoogleMedium49$1.25$5.00$0.0250Moderate (6-20s)Standard (2/4)General enterprise analysis
Gemini 3.8 FlashGoogleHigh48$0.10$0.40$0.0075Deliberate (30-150s)Deep (3/4)High-volume complex analysis
Gemini 3.8 FlashGoogleLow35$0.10$0.40$0.0009Fast (1-5s)Check (1/4)High-speed reasoning verification
Gemini 3.8 FlashGoogleMax50$0.10$0.40$0.0120Deliberate (30-150s)Exhaustive (4/4)Budget frontier reasoning
Gemini 3.8 FlashGoogleMedium44$0.10$0.40$0.0022Moderate (6-20s)Standard (2/4)General workhorse coding and reasoning
GPT-5.6 LunaOpenAIExtra High35$0.20$1.20$0.1200Deliberate (30-150s)Deep (3/4)Hard problems, higher accuracy
GPT-5.6 LunaOpenAIHigh33$0.20$1.20$0.0900Deliberate (30-150s)Deep (3/4)Complex but non-frontier tasks
GPT-5.6 LunaOpenAILow22$0.20$1.20$0.0400Fast (1-5s)Check (1/4)Simple, high-volume tasks
GPT-5.6 LunaOpenAIMax38$0.20$1.20$0.1800Deliberate (30-150s)Deep (3/4)Toughest tasks on a budget
GPT-5.6 LunaOpenAIMedium26$0.20$1.20$0.0600Moderate (6-20s)Standard (2/4)General everyday work
GPT-5.6 SolOpenAIExtra High44$4.00$20.00$1.4300Deliberate (30-150s)Deep (3/4)Very challenging tasks
GPT-5.6 SolOpenAIHigh42$4.00$20.00$0.8100Deliberate (30-150s)Deep (3/4)Hard problems
GPT-5.6 SolOpenAILow34$4.00$20.00$0.2300Moderate (6-20s)Check (1/4)Everyday tasks with stronger results
GPT-5.6 SolOpenAIMax47$4.00$20.00$1.9900Deliberate (30-150s)Exhaustive (4/4)Frontier-level tasks
GPT-5.6 SolOpenAIMedium39$4.00$20.00$0.5000Moderate (6-20s)Standard (2/4)Complex everyday work
GPT-5.6 TerraOpenAIExtra High38$2.00$12.00$0.9300Deliberate (30-150s)Deep (3/4)Challenging, high-accuracy tasks
GPT-5.6 TerraOpenAIHigh34$2.00$12.00$0.5200Deliberate (30-150s)Deep (3/4)Complex tasks
GPT-5.6 TerraOpenAILow28$2.00$12.00$0.1200Moderate (6-20s)Check (1/4)Simple tasks at higher quality than Luna
GPT-5.6 TerraOpenAIMax42$2.00$12.00$1.4000Deliberate (30-150s)Exhaustive (4/4)Maximum capability when needed
GPT-5.6 TerraOpenAIMedium33$2.00$12.00$0.3300Moderate (6-20s)Standard (2/4)General work with more depth
GPT-6 AstraOpenAIExtra High53$10.00$50.00$3.2600Deliberate (30-150s)Deep (3/4)Near-maximum capability
GPT-6 AstraOpenAIHigh51$10.00$50.00$2.7100Deliberate (30-150s)Deep (3/4)Very complex tasks
GPT-6 AstraOpenAILow46$10.00$50.00$0.8200Moderate (6-20s)Check (1/4)High intelligence at reasonable cost
GPT-6 AstraOpenAIMax53$10.00$50.00$3.2600Deliberate (30-150s)Exhaustive (4/4)Hardest possible tasks
GPT-6 AstraOpenAIMedium50$10.00$50.00$2.1700Moderate (6-20s)Standard (2/4)Challenging, nuanced tasks

References

Official provider documentation, technical releases, and benchmark evaluation cards. Every listed provider model and pricing tier is verified against primary documentation.

  1. OpenAI Reasoning Models and PricingOpenAI · API specifications and reported evaluations

    Official pricing, reasoning effort parameter definitions, token rates, and evaluation benchmarks for Astra, Sol, Terra, and Luna model tiers.

    Published/update: 2026-09-03; accessed 2026-09-19.
  2. Claude 3.7 Sonnet and Extended Thinking SpecificationsAnthropic · Provider release, pricing, and reported evaluations

    Hybrid reasoning architecture details, budget token controls, output token rates, SWE-bench and AIME evaluation scaling across thinking depths.

    Published/update: 2025-02-24; accessed 2026-09-19.
  3. Gemini 3.8 Flash and Gemini 3.7 Flash and Pro DocumentationGoogle · Provider API metadata and serving rates

    Thinking budget parameter, latency profiles, token pricing tiers, and mathematical reasoning scores across Flash and Pro variants.

    Published/update: 2026-09-17; accessed 2026-09-19.
  4. Frontier Model Intelligence and Quality EvaluationsIndependent benchmarks and provider cards · Methodology and context

    Normalized 0-100 intelligence indexing methodology combining STEM benchmarks, reasoning evaluations, and speed-latency trade-offs.

    Published/update: 2026-09-08; accessed 2026-09-19.