Compare selected current API and open-weight model releases with benchmark evidence, price, and serving metadata.
Snapshot accessed 7 selected model/version rows · release dates retain source precision; a page update is not treated as a release date.
SELECTED MODELS7
Named releases in this dated snapshot.
API AVAILABLE7
Rows with a retained hosted endpoint reference.
OPEN-WEIGHT4
Rows with weights or self-hosting evidence.
REPORTED SCORES4
Models with provider/model-card percentages.
Observations
These findings describe the full dated snapshot. Filtering records below does not change the analysis.
Evidence is incomplete. 4 of 7 models have retained numeric evaluation results, and 4 have both token prices. Missing evidence is unknown, not zero performance; coverage reflects our retained sources.
Listed input prices range from $0.50 to $10.00. Across 4 priced models, per million input tokens. Mistral Large 3 is the lower endpoint; GPT-6 Astra is an upper endpoint. Output costs are separate, and price alone does not establish quality or total task cost.
Scoring rules change the result by 36.2 points. Claude Fable 5.1 reports 77.9% partial completion and 41.7% strict completion on OSWorld 2.0. These are different scoring rules, not a change in model ability over time. Source: anthropic-fable51 in References.
PROVIDER / MODEL-CARD RESULTS · PERCENT
Reported benchmark results
one model at a time
Each bar is a separate provider-reported evaluation. Tasks, scoring definitions, tools, and settings differ. Compare the partial and strict OSWorld scores for Claude Fable 5.1 to see why a score needs its definition. This is not an independent comparison of models.
View all reported evaluation data
Accessible data table, reported benchmark percentages
Model
Provider
Benchmark/version
Score
Evaluation setting
Source IDs
GLM-5.3-Flash
Z.ai
Terminal-Bench 2.1
84.3%
Temperature 1; max 65,536; 6-hour limit
glm53-card
GLM-5.3-Flash
Z.ai
DeepSWE
63.4%
Temperature .95; max 400,000; 6-hour limit
glm53-card
GLM-5.3-Flash
Z.ai
Humanity's Last Exam
55.3%
Temperature 1; max 163,840; 300k context
glm53-card
GPT-6 Astra
OpenAI
Terminal-Bench 4.0
57.9%
OpenAI evaluation; provider setting retained in launch report
openai-api, openai-gpt6-launch
GPT-6 Astra
OpenAI
GPQA Diamond
96.0%
OpenAI evaluation; provider setting retained in launch report
openai-api, openai-gpt6-launch
Claude Fable 5.1
Anthropic
Terminal-Bench-Science 0.1
52.6%
Production safeguards enabled; Anthropic setup
anthropic-fable51
Claude Fable 5.1
Anthropic
Terminal-Bench 4.0
55.8%
Production safeguards enabled; Anthropic setup
anthropic-fable51
Claude Fable 5.1
Anthropic
Humanity's Last Exam (no tools)
60.9%
Production safeguards enabled; no tools
anthropic-fable51
Claude Fable 5.1
Anthropic
Humanity's Last Exam (with tools)
65.0%
Production safeguards enabled; with tools
anthropic-fable51
Claude Fable 5.1
Anthropic
CursorBench 3.2.0
73.4%
Max effort; Anthropic report
anthropic-fable51
Claude Fable 5.1
Anthropic
OSWorld 2.0 (partial)
77.9%
August 2026 task release; partial completion; production safeguards enabled
anthropic-fable51
Claude Fable 5.1
Anthropic
OSWorld 2.0 (strict)
41.7%
August 2026 task release; strict completion; production safeguards enabled
anthropic-fable51
Qwen3.8 2.4T A95B
Qwen
Terminal-Bench 2.1
86.6%
Qwen evaluation; harness and run settings retained in model card
qwen38-card
Qwen3.8 2.4T A95B
Qwen
DeepSWE
56.6%
Qwen evaluation; harness and run settings retained in model card
qwen38-card
Qwen3.8 2.4T A95B
Qwen
SWE-bench Pro
67.7%
Qwen evaluation; harness and run settings retained in model card
qwen38-card
Qwen3.8 2.4T A95B
Qwen
PaperBench
93.0%
Qwen evaluation; harness and run settings retained in model card
qwen38-card
LISTED TOKEN PRICE · USD / 1M TOKENS
Input and output price
metadata, not quality
The chart uses the first listed USD tier for rows with both prices. Google Gemini 3.8 Flash has an introductory tier through 2026-12-31 and a standard tier from 2027-01-01; the full transition stays in the record table.
Accessible data table, listed price in USD per one million tokens
Model
Provider
Input
Output
Gemini 3.8 Flash
Google
$0.75
$3.75
GPT-6 Astra
OpenAI
$10.00
$50.00
Claude Fable 5.1
Anthropic
$10.00
$50.00
Mistral Large 3
Mistral AI
$0.50
$1.50
About the data
Coverage is selected, not exhaustive. The snapshot prioritizes current named releases in the retained provider sources; Mistral Large 3 is included as a dated open-weight large-model reference and is not presented as a claim about every newer Mistral release. “Open-weight” means weights or self-hosting evidence were retained; it does not automatically mean an OSI-approved open-source license. The performance chart shows separate evaluations for one selected model. Benchmark definitions, settings, and harnesses differ, so it does not claim a universal model ranking. Prices, latency, and throughput are serving metadata and are not quality scores.
3 row(s) have no numeric benchmark retained from the listed sources. Missing values stay visible as “Not published in retained source”.
Tentative, evaluation design changes the result. Partial-credit scoring rewards progress that strict completion does not. That definition is consistent with the OSWorld score gap, but this snapshot cannot isolate the effects of tools, safeguards, or run settings. A matched rerun with the same model and harness would test those effects.
Hypothesis, token prices reflect different serving choices. Infrastructure costs, introductory pricing, and product positioning could contribute to the spread. The price list alone does not reveal provider costs or performance per dollar. Test a fixed task set with measured tokens, successful outcomes, latency, and the applicable billing tier before estimating value.
Open-weight; license is not OSI-identified in retained source
Qwen3.8 Max custom license
2.4T total / 95B active
262,144 native / 1,010,000 extensible tokens
Not published in retained source
Not published in retained source
Not published in retained source
Not published in retained source
Terminal-Bench 2.1: 86.6% (Model-card reported; setting: Qwen evaluation; harness and run settings retained in model card; source: qwen38-card); DeepSWE: 56.6% (Model-card reported; setting: Qwen evaluation; harness and run settings retained in model card; source: qwen38-card); SWE-bench Pro: 67.7% (Model-card reported; setting: Qwen evaluation; harness and run settings retained in model card; source: qwen38-card); PaperBench: 93.0% (Model-card reported; setting: Qwen evaluation; harness and run settings retained in model card; source: qwen38-card)
Rows retain the normalized source fields. Download JSON.
References
These are the substantive provider/model-card sources used for this dated page. The record table maps benchmark settings and source IDs to the retained evidence.