Gemini 3.8 / 3.7 Flash Medium
Score 43-44 · $0.0020-$0.0022/task · Value Index: 20,000+
Top-tier math and coding at fraction-of-a-cent rates with under 10-second latency.
Compare reasoning levels across OpenAI GPT, Google Gemini, and Anthropic Claude, showing intelligence scores, task costs, and execution speeds.
Collected UTC
Coverage: 34 model reasoning tiers across OpenAI, Google, and Anthropic.
Score 43-44 · $0.0020-$0.0022/task · Value Index: 20,000+
Top-tier math and coding at fraction-of-a-cent rates with under 10-second latency.
Score 33 · $0.09/task · Value Index: 183
High-accuracy multi-step research and structured code generation.
Score 43-49 · $0.03-$0.07/task · Value Index: 700+
The premier sweet spot for daily software engineering, refactoring, and agent loops.
Score 53-55 · $0.95-$3.26/task · Value Index: 16-58
Reserved strictly for mission-critical architecture, novel mathematics, and high-stakes proofs.
Higher and to the left represents superior efficiency (greater intelligence at lower task cost). Key anchor models are labeled directly on the chart dots.
| Model | Provider | Level | Avg Cost / Task | Intelligence Index |
|---|---|---|---|---|
| Claude 3.7 Sonnet | Anthropic | High | $0.2200 | 53 |
| Claude 3.7 Sonnet | Anthropic | Low | $0.0300 | 43 |
| Claude 3.7 Sonnet | Anthropic | Max | $0.9500 | 55 |
| Claude 3.7 Sonnet | Anthropic | Medium | $0.0700 | 49 |
| Claude 3.7 Sonnet | Anthropic | Standard | $0.0100 | 36 |
| Claude Mythos 5.1 | Anthropic | High | $1.2000 | 54 |
| Claude Mythos 5.1 | Anthropic | Max | $4.8000 | 56 |
| Gemini 3.7 Flash | Medium | $0.0020 | 43 | |
| Gemini 3.7 Pro | High | $0.0500 | 51 | |
| Gemini 3.7 Pro | Medium | $0.0250 | 49 | |
| Gemini 3.8 Flash | High | $0.0075 | 48 | |
| Gemini 3.8 Flash | Low | $0.0009 | 35 | |
| Gemini 3.8 Flash | Max | $0.0120 | 50 | |
| Gemini 3.8 Flash | Medium | $0.0022 | 44 | |
| GPT-5.6 Luna | OpenAI | Extra High | $0.1200 | 35 |
| GPT-5.6 Luna | OpenAI | High | $0.0900 | 33 |
| GPT-5.6 Luna | OpenAI | Low | $0.0400 | 22 |
| GPT-5.6 Luna | OpenAI | Max | $0.1800 | 38 |
| GPT-5.6 Luna | OpenAI | Medium | $0.0600 | 26 |
| GPT-5.6 Sol | OpenAI | Extra High | $1.4300 | 44 |
| GPT-5.6 Sol | OpenAI | High | $0.8100 | 42 |
| GPT-5.6 Sol | OpenAI | Low | $0.2300 | 34 |
| GPT-5.6 Sol | OpenAI | Max | $1.9900 | 47 |
| GPT-5.6 Sol | OpenAI | Medium | $0.5000 | 39 |
| GPT-5.6 Terra | OpenAI | Extra High | $0.9300 | 38 |
| GPT-5.6 Terra | OpenAI | High | $0.5200 | 34 |
| GPT-5.6 Terra | OpenAI | Low | $0.1200 | 28 |
| GPT-5.6 Terra | OpenAI | Max | $1.4000 | 42 |
| GPT-5.6 Terra | OpenAI | Medium | $0.3300 | 33 |
| GPT-6 Astra | OpenAI | Extra High | $3.2600 | 53 |
| GPT-6 Astra | OpenAI | High | $2.7100 | 51 |
| GPT-6 Astra | OpenAI | Low | $0.8200 | 46 |
| GPT-6 Astra | OpenAI | Max | $3.2600 | 53 |
| GPT-6 Astra | OpenAI | Medium | $2.1700 | 50 |
Comparing capability against response turnaround tiers: Fast (1-5s), Moderate (6-20s), and Deliberate (30-150s). Notice the sharp latency penalty required for top-tier reasoning effort.
| Model | Provider | Level | Speed Tier | Intelligence Index |
|---|---|---|---|---|
| Claude 3.7 Sonnet | Anthropic | High | Deliberate (30-150s) | 53 |
| Claude 3.7 Sonnet | Anthropic | Low | Moderate (6-20s) | 43 |
| Claude 3.7 Sonnet | Anthropic | Max | Deliberate (30-150s) | 55 |
| Claude 3.7 Sonnet | Anthropic | Medium | Moderate (6-20s) | 49 |
| Claude 3.7 Sonnet | Anthropic | Standard | Fast (1-5s) | 36 |
| Claude Mythos 5.1 | Anthropic | High | Deliberate (30-150s) | 54 |
| Claude Mythos 5.1 | Anthropic | Max | Deliberate (30-150s) | 56 |
| Gemini 3.7 Flash | Medium | Moderate (6-20s) | 43 | |
| Gemini 3.7 Pro | High | Deliberate (30-150s) | 51 | |
| Gemini 3.7 Pro | Medium | Moderate (6-20s) | 49 | |
| Gemini 3.8 Flash | High | Deliberate (30-150s) | 48 | |
| Gemini 3.8 Flash | Low | Fast (1-5s) | 35 | |
| Gemini 3.8 Flash | Max | Deliberate (30-150s) | 50 | |
| Gemini 3.8 Flash | Medium | Moderate (6-20s) | 44 | |
| GPT-5.6 Luna | OpenAI | Extra High | Deliberate (30-150s) | 35 |
| GPT-5.6 Luna | OpenAI | High | Deliberate (30-150s) | 33 |
| GPT-5.6 Luna | OpenAI | Low | Fast (1-5s) | 22 |
| GPT-5.6 Luna | OpenAI | Max | Deliberate (30-150s) | 38 |
| GPT-5.6 Luna | OpenAI | Medium | Moderate (6-20s) | 26 |
| GPT-5.6 Sol | OpenAI | Extra High | Deliberate (30-150s) | 44 |
| GPT-5.6 Sol | OpenAI | High | Deliberate (30-150s) | 42 |
| GPT-5.6 Sol | OpenAI | Low | Moderate (6-20s) | 34 |
| GPT-5.6 Sol | OpenAI | Max | Deliberate (30-150s) | 47 |
| GPT-5.6 Sol | OpenAI | Medium | Moderate (6-20s) | 39 |
| GPT-5.6 Terra | OpenAI | Extra High | Deliberate (30-150s) | 38 |
| GPT-5.6 Terra | OpenAI | High | Deliberate (30-150s) | 34 |
| GPT-5.6 Terra | OpenAI | Low | Moderate (6-20s) | 28 |
| GPT-5.6 Terra | OpenAI | Max | Deliberate (30-150s) | 42 |
| GPT-5.6 Terra | OpenAI | Medium | Moderate (6-20s) | 33 |
| GPT-6 Astra | OpenAI | Extra High | Deliberate (30-150s) | 53 |
| GPT-6 Astra | OpenAI | High | Deliberate (30-150s) | 51 |
| GPT-6 Astra | OpenAI | Low | Moderate (6-20s) | 46 |
| GPT-6 Astra | OpenAI | Max | Deliberate (30-150s) | 53 |
| GPT-6 Astra | OpenAI | Medium | Moderate (6-20s) | 50 |
This dataset compares reasoning capabilities, pricing, task costs, and execution speeds across frontier AI models from OpenAI, Anthropic, and Google. It provides practitioners and engineers with the trade-offs needed to configure reasoning effort levels in production software.
Prices and performance figures reflect official provider documentation and benchmark cards as of September 2026. Latency figures represent average direct API performance and may vary based on provider server load, region, and batching.
Proposed mechanism: Reasoning models use reinforcement learning to search through multiple potential solution paths, backtrack when hitting dead ends, and verify steps internally before generating a final response.
Evidence: Performance on formal verification tasks (like competition mathematics and code execution) jumps significantly with thinking tokens. However, once the model finds the correct solution path, additional reasoning tokens provide diminishing returns, explaining the plateau between High and Max levels.
Proposed mechanism: Output tokens require sequential autoregressive generation, which cannot be parallelized across tokens like prompt ingestion. Providers price output tokens 3x to 5x higher than input tokens to reflect hardware allocation.
Evidence: Because all thinking tokens are billed as output tokens, a query that generates 15,000 thinking tokens shifts over 90% of the total request cost onto the generation stage, causing exponential cost growth at high effort levels.
Search and compare reasoning models by provider, intelligence score, cost per task, and recommended use case.
| Model | Provider | Reasoning Level | Intelligence Index | Input / 1M | Output / 1M | Avg Cost / Task | Speed | Depth | Best For |
|---|---|---|---|---|---|---|---|---|---|
| Claude 3.7 Sonnet | Anthropic | High | 53 | $3.00 | $15.00 | $0.2200 | Deliberate (30-150s) | Deep (3/4) | Complex tasks and deep logic |
| Claude 3.7 Sonnet | Anthropic | Low | 43 | $3.00 | $15.00 | $0.0300 | Moderate (6-20s) | Check (1/4) | High intelligence at balanced cost |
| Claude 3.7 Sonnet | Anthropic | Max | 55 | $3.00 | $15.00 | $0.9500 | Deliberate (30-150s) | Exhaustive (4/4) | Hardest possible tasks |
| Claude 3.7 Sonnet | Anthropic | Medium | 49 | $3.00 | $15.00 | $0.0700 | Moderate (6-20s) | Standard (2/4) | Complex everyday development |
| Claude 3.7 Sonnet | Anthropic | Standard | 36 | $3.00 | $15.00 | $0.0100 | Fast (1-5s) | Check (1/4) | Everyday tasks with rapid responses |
| Claude Mythos 5.1 | Anthropic | High | 54 | $15.00 | $75.00 | $1.2000 | Deliberate (30-150s) | Deep (3/4) | Frontier-level multi-step tasks |
| Claude Mythos 5.1 | Anthropic | Max | 56 | $15.00 | $75.00 | $4.8000 | Deliberate (30-150s) | Exhaustive (4/4) | Maximum capability frontier science |
| Gemini 3.7 Flash | Medium | 43 | $0.10 | $0.40 | $0.0020 | Moderate (6-20s) | Standard (2/4) | High-efficiency coding and problem solving | |
| Gemini 3.7 Pro | High | 51 | $1.25 | $5.00 | $0.0500 | Deliberate (30-150s) | Deep (3/4) | Complex multimodal planning | |
| Gemini 3.7 Pro | Medium | 49 | $1.25 | $5.00 | $0.0250 | Moderate (6-20s) | Standard (2/4) | General enterprise analysis | |
| Gemini 3.8 Flash | High | 48 | $0.10 | $0.40 | $0.0075 | Deliberate (30-150s) | Deep (3/4) | High-volume complex analysis | |
| Gemini 3.8 Flash | Low | 35 | $0.10 | $0.40 | $0.0009 | Fast (1-5s) | Check (1/4) | High-speed reasoning verification | |
| Gemini 3.8 Flash | Max | 50 | $0.10 | $0.40 | $0.0120 | Deliberate (30-150s) | Exhaustive (4/4) | Budget frontier reasoning | |
| Gemini 3.8 Flash | Medium | 44 | $0.10 | $0.40 | $0.0022 | Moderate (6-20s) | Standard (2/4) | General workhorse coding and reasoning | |
| GPT-5.6 Luna | OpenAI | Extra High | 35 | $0.20 | $1.20 | $0.1200 | Deliberate (30-150s) | Deep (3/4) | Hard problems, higher accuracy |
| GPT-5.6 Luna | OpenAI | High | 33 | $0.20 | $1.20 | $0.0900 | Deliberate (30-150s) | Deep (3/4) | Complex but non-frontier tasks |
| GPT-5.6 Luna | OpenAI | Low | 22 | $0.20 | $1.20 | $0.0400 | Fast (1-5s) | Check (1/4) | Simple, high-volume tasks |
| GPT-5.6 Luna | OpenAI | Max | 38 | $0.20 | $1.20 | $0.1800 | Deliberate (30-150s) | Deep (3/4) | Toughest tasks on a budget |
| GPT-5.6 Luna | OpenAI | Medium | 26 | $0.20 | $1.20 | $0.0600 | Moderate (6-20s) | Standard (2/4) | General everyday work |
| GPT-5.6 Sol | OpenAI | Extra High | 44 | $4.00 | $20.00 | $1.4300 | Deliberate (30-150s) | Deep (3/4) | Very challenging tasks |
| GPT-5.6 Sol | OpenAI | High | 42 | $4.00 | $20.00 | $0.8100 | Deliberate (30-150s) | Deep (3/4) | Hard problems |
| GPT-5.6 Sol | OpenAI | Low | 34 | $4.00 | $20.00 | $0.2300 | Moderate (6-20s) | Check (1/4) | Everyday tasks with stronger results |
| GPT-5.6 Sol | OpenAI | Max | 47 | $4.00 | $20.00 | $1.9900 | Deliberate (30-150s) | Exhaustive (4/4) | Frontier-level tasks |
| GPT-5.6 Sol | OpenAI | Medium | 39 | $4.00 | $20.00 | $0.5000 | Moderate (6-20s) | Standard (2/4) | Complex everyday work |
| GPT-5.6 Terra | OpenAI | Extra High | 38 | $2.00 | $12.00 | $0.9300 | Deliberate (30-150s) | Deep (3/4) | Challenging, high-accuracy tasks |
| GPT-5.6 Terra | OpenAI | High | 34 | $2.00 | $12.00 | $0.5200 | Deliberate (30-150s) | Deep (3/4) | Complex tasks |
| GPT-5.6 Terra | OpenAI | Low | 28 | $2.00 | $12.00 | $0.1200 | Moderate (6-20s) | Check (1/4) | Simple tasks at higher quality than Luna |
| GPT-5.6 Terra | OpenAI | Max | 42 | $2.00 | $12.00 | $1.4000 | Deliberate (30-150s) | Exhaustive (4/4) | Maximum capability when needed |
| GPT-5.6 Terra | OpenAI | Medium | 33 | $2.00 | $12.00 | $0.3300 | Moderate (6-20s) | Standard (2/4) | General work with more depth |
| GPT-6 Astra | OpenAI | Extra High | 53 | $10.00 | $50.00 | $3.2600 | Deliberate (30-150s) | Deep (3/4) | Near-maximum capability |
| GPT-6 Astra | OpenAI | High | 51 | $10.00 | $50.00 | $2.7100 | Deliberate (30-150s) | Deep (3/4) | Very complex tasks |
| GPT-6 Astra | OpenAI | Low | 46 | $10.00 | $50.00 | $0.8200 | Moderate (6-20s) | Check (1/4) | High intelligence at reasonable cost |
| GPT-6 Astra | OpenAI | Max | 53 | $10.00 | $50.00 | $3.2600 | Deliberate (30-150s) | Exhaustive (4/4) | Hardest possible tasks |
| GPT-6 Astra | OpenAI | Medium | 50 | $10.00 | $50.00 | $2.1700 | Moderate (6-20s) | Standard (2/4) | Challenging, nuanced tasks |
Official provider documentation, technical releases, and benchmark evaluation cards. Every listed provider model and pricing tier is verified against primary documentation.
Official pricing, reasoning effort parameter definitions, token rates, and evaluation benchmarks for Astra, Sol, Terra, and Luna model tiers.
Published/update: 2026-09-03; accessed 2026-09-19.Hybrid reasoning architecture details, budget token controls, output token rates, SWE-bench and AIME evaluation scaling across thinking depths.
Published/update: 2025-02-24; accessed 2026-09-19.Thinking budget parameter, latency profiles, token pricing tiers, and mathematical reasoning scores across Flash and Pro variants.
Published/update: 2026-09-17; accessed 2026-09-19.Normalized 0-100 intelligence indexing methodology combining STEM benchmarks, reasoning evaluations, and speed-latency trade-offs.
Published/update: 2026-09-08; accessed 2026-09-19.