Back to information library
Artificial intelligence / DATED RESEARCH SNAPSHOT

Which AI models are available now?

Compare selected current API and open-weight model releases with benchmark evidence, price, and serving metadata.

Snapshot accessed
7 selected model/version rows · release dates retain source precision; a page update is not treated as a release date.

SELECTED MODELS7

Named releases in this dated snapshot.

API AVAILABLE7

Rows with a retained hosted endpoint reference.

OPEN-WEIGHT4

Rows with weights or self-hosting evidence.

REPORTED SCORES4

Models with provider/model-card percentages.

Observations

These findings describe the full dated snapshot. Filtering records below does not change the analysis.

PROVIDER / MODEL-CARD RESULTS · PERCENT

Reported benchmark results

one model at a time

Each bar is a separate provider-reported evaluation. Tasks, scoring definitions, tools, and settings differ. Compare the partial and strict OSWorld scores for Claude Fable 5.1 to see why a score needs its definition. This is not an independent comparison of models.

View all reported evaluation data
Accessible data table, reported benchmark percentages
ModelProviderBenchmark/versionScoreEvaluation settingSource IDs
GLM-5.3-FlashZ.aiTerminal-Bench 2.184.3%Temperature 1; max 65,536; 6-hour limitglm53-card
GLM-5.3-FlashZ.aiDeepSWE63.4%Temperature .95; max 400,000; 6-hour limitglm53-card
GLM-5.3-FlashZ.aiHumanity's Last Exam55.3%Temperature 1; max 163,840; 300k contextglm53-card
GPT-6 AstraOpenAITerminal-Bench 4.057.9%OpenAI evaluation; provider setting retained in launch reportopenai-api, openai-gpt6-launch
GPT-6 AstraOpenAIGPQA Diamond96.0%OpenAI evaluation; provider setting retained in launch reportopenai-api, openai-gpt6-launch
Claude Fable 5.1AnthropicTerminal-Bench-Science 0.152.6%Production safeguards enabled; Anthropic setupanthropic-fable51
Claude Fable 5.1AnthropicTerminal-Bench 4.055.8%Production safeguards enabled; Anthropic setupanthropic-fable51
Claude Fable 5.1AnthropicHumanity's Last Exam (no tools)60.9%Production safeguards enabled; no toolsanthropic-fable51
Claude Fable 5.1AnthropicHumanity's Last Exam (with tools)65.0%Production safeguards enabled; with toolsanthropic-fable51
Claude Fable 5.1AnthropicCursorBench 3.2.073.4%Max effort; Anthropic reportanthropic-fable51
Claude Fable 5.1AnthropicOSWorld 2.0 (partial)77.9%August 2026 task release; partial completion; production safeguards enabledanthropic-fable51
Claude Fable 5.1AnthropicOSWorld 2.0 (strict)41.7%August 2026 task release; strict completion; production safeguards enabledanthropic-fable51
Qwen3.8 2.4T A95BQwenTerminal-Bench 2.186.6%Qwen evaluation; harness and run settings retained in model cardqwen38-card
Qwen3.8 2.4T A95BQwenDeepSWE56.6%Qwen evaluation; harness and run settings retained in model cardqwen38-card
Qwen3.8 2.4T A95BQwenSWE-bench Pro67.7%Qwen evaluation; harness and run settings retained in model cardqwen38-card
Qwen3.8 2.4T A95BQwenPaperBench93.0%Qwen evaluation; harness and run settings retained in model cardqwen38-card
LISTED TOKEN PRICE · USD / 1M TOKENS

Input and output price

metadata, not quality

The chart uses the first listed USD tier for rows with both prices. Google Gemini 3.8 Flash has an introductory tier through 2026-12-31 and a standard tier from 2027-01-01; the full transition stays in the record table.

Accessible data table, listed price in USD per one million tokens
ModelProviderInputOutput
Gemini 3.8 FlashGoogle$0.75$3.75
GPT-6 AstraOpenAI$10.00$50.00
Claude Fable 5.1Anthropic$10.00$50.00
Mistral Large 3Mistral AI$0.50$1.50

About the data

Coverage is selected, not exhaustive. The snapshot prioritizes current named releases in the retained provider sources; Mistral Large 3 is included as a dated open-weight large-model reference and is not presented as a claim about every newer Mistral release. “Open-weight” means weights or self-hosting evidence were retained; it does not automatically mean an OSI-approved open-source license. The performance chart shows separate evaluations for one selected model. Benchmark definitions, settings, and harnesses differ, so it does not claim a universal model ranking. Prices, latency, and throughput are serving metadata and are not quality scores.

3 row(s) have no numeric benchmark retained from the listed sources. Missing values stay visible as “Not published in retained source”.

Source ledger: provenance download · research/ai-models-snapshot-2026-09-19.json · methodology

Possible explanations

Tentative, evaluation design changes the result. Partial-credit scoring rewards progress that strict completion does not. That definition is consistent with the OSWorld score gap, but this snapshot cannot isolate the effects of tools, safeguards, or run settings. A matched rerun with the same model and harness would test those effects.

Hypothesis, token prices reflect different serving choices. Infrastructure costs, introductory pricing, and product positioning could contribute to the spread. The price list alone does not reveal provider costs or performance per dollar. Test a fixed task set with measured tokens, successful outcomes, latency, and the applicable billing tier before estimating value.

Explore the records

Download all records (CSV)

7 records

ModelProviderVersionReleaseAPI / servingOpennessLicenseParametersContextInput priceOutput priceLatencyThroughputReported benchmark evidenceEvidence
GLM-5.3-FlashZ.aiGLM-5.3-Flash2026-09-19Z.ai API; self-hostable via vLLM/SGLang/TransformersOpen-weight and model materials under MITMIT320B total / 18B active300,000 tokens in retained evaluation settingNot published in retained sourceNot published in retained sourceNot published in retained sourceNot published in retained sourceTerminal-Bench 2.1: 84.3% (Model-card reported; setting: Temperature 1; max 65,536; 6-hour limit; source: glm53-card); DeepSWE: 63.4% (Model-card reported; setting: Temperature .95; max 400,000; 6-hour limit; source: glm53-card); Humanity's Last Exam: 55.3% (Model-card reported; setting: Temperature 1; max 163,840; 300k context; source: glm53-card)Evidence
Gemini 3.8 FlashGooglegemini-3.8-flash2026-09-17 page update; release date not statedGemini APIProprietary API modelGoogle API termsNot published in retained source1,048,576 input / 65,536 output tokens$0.75 / 1M tokens through 2026-12-31; $1.50 from 2027-01-01$3.75 / 1M tokens through 2026-12-31; $7.50 from 2027-01-01Not published in retained sourceNot published in retained sourceNot published in retained sourceEvidence
DeepSeek V4.1-FlashDeepSeekdeepseek-flash2026-09-10DeepSeek API (`deepseek-flash`); provider links weightsOpen-weight status indicated by provider; license not verified in retained sourceNot published in retained source552B total / 8B input-active / 16B output-activeNot published in retained sourceNot published in retained sourceNot published in retained sourceNot published in retained sourceNot published in retained sourceNot published in retained sourceEvidence
GPT-6 AstraOpenAIgpt-6-astra2026-09-03OpenAI API; Azure; Amazon BedrockProprietary API modelProvider termsNot published in retained source1,050,000 input / 131,072 output tokens$10.00 / 1M tokens$50.00 / 1M tokensNot published in retained sourceNot published in retained sourceTerminal-Bench 4.0: 57.9% (Provider reported; setting: OpenAI evaluation; provider setting retained in launch report; source: openai-api, openai-gpt6-launch); GPQA Diamond: 96.0% (Provider reported; setting: OpenAI evaluation; provider setting retained in launch report; source: openai-api, openai-gpt6-launch)Evidence
Claude Fable 5.1Anthropicclaude-fable-5-12026-09-01Claude API; Amazon Web Services; Google Cloud; Microsoft AzureProprietary API modelProvider termsNot published in retained sourceNot published in retained source$10.00 / 1M tokens$50.00 / 1M tokensNot published in retained sourceNot published in retained sourceTerminal-Bench-Science 0.1: 52.6% (Provider reported; setting: Production safeguards enabled; Anthropic setup; source: anthropic-fable51); Terminal-Bench 4.0: 55.8% (Provider reported; setting: Production safeguards enabled; Anthropic setup; source: anthropic-fable51); Humanity's Last Exam (no tools): 60.9% (Provider reported; setting: Production safeguards enabled; no tools; source: anthropic-fable51); Humanity's Last Exam (with tools): 65.0% (Provider reported; setting: Production safeguards enabled; with tools; source: anthropic-fable51); CursorBench 3.2.0: 73.4% (Provider reported; setting: Max effort; Anthropic report; source: anthropic-fable51); OSWorld 2.0 (partial): 77.9% (Provider reported; setting: August 2026 task release; partial completion; production safeguards enabled; source: anthropic-fable51); OSWorld 2.0 (strict): 41.7% (Provider reported; setting: August 2026 task release; strict completion; production safeguards enabled; source: anthropic-fable51)Evidence
Qwen3.8 2.4T A95BQwenQwen3.8-2.4T-A95B2026-08-12 page update; release date not statedQwen Cloud API; self-hostableOpen-weight; license is not OSI-identified in retained sourceQwen3.8 Max custom license2.4T total / 95B active262,144 native / 1,010,000 extensible tokensNot published in retained sourceNot published in retained sourceNot published in retained sourceNot published in retained sourceTerminal-Bench 2.1: 86.6% (Model-card reported; setting: Qwen evaluation; harness and run settings retained in model card; source: qwen38-card); DeepSWE: 56.6% (Model-card reported; setting: Qwen evaluation; harness and run settings retained in model card; source: qwen38-card); SWE-bench Pro: 67.7% (Model-card reported; setting: Qwen evaluation; harness and run settings retained in model card; source: qwen38-card); PaperBench: 93.0% (Model-card reported; setting: Qwen evaluation; harness and run settings retained in model card; source: qwen38-card)Evidence
Mistral Large 3Mistral AImistral-large-32025-12-02Mistral API; Amazon Bedrock; Microsoft AzureOpen-weightApache 2.0675B total / 41B active256,000 tokens$0.50 / 1M tokens$1.50 / 1M tokensNot published in retained sourceNot published in retained sourceNot published in retained sourceEvidence

Rows retain the normalized source fields. Download JSON.

References

These are the substantive provider/model-card sources used for this dated page. The record table maps benchmark settings and source IDs to the retained evidence.

  1. GPT-6 Astra model documentation, OpenAI; Provider API metadata. Published/update: 2026-09-03; accessed 2026-09-19.
  2. Introducing GPT-6 Astra, OpenAI; Provider release and reported evaluations. Published/update: 2026-09-03; accessed 2026-09-19.
  3. Introducing Claude Fable 5.1 and Claude Mythos 5.1, Anthropic; Provider release, pricing, and reported evaluations. Published/update: 2026-09-01; accessed 2026-09-19.
  4. Gemini 3.8 Flash model documentation, Google; Provider API metadata. Published/update: 2026-09-17 page update; accessed 2026-09-19.
  5. Gemini API latest-model guide and pricing, Google; Provider price and release context. Published/update: 2026-09-17 page update; accessed 2026-09-19.
  6. Qwen3.8 2.4T A95B model card, Qwen / Hugging Face; Official model card and reported evaluations. Published/update: 2026-08-12 page update; accessed 2026-09-19.
  7. GLM-5.3-Flash model card, Z.ai / Hugging Face; Official model card and reported evaluations. Published/update: 2026-09-19; accessed 2026-09-19.
  8. Mistral 3, Mistral AI; Provider release and license. Published/update: 2025-12-02; accessed 2026-09-19.
  9. Inference pricing, Mistral AI; Provider API pricing. Published/update: 2026-09; accessed 2026-09-19.
  10. DeepSeek V4.1-Flash, DeepSeek; Provider release and API identifier. Published/update: 2026-09-10; accessed 2026-09-19.

Visualization package credits