AI Model Roundup: Sep 7-13, 2026
This week was quieter on new launches than the Sep 1-6 stretch, but it produced real news. Mercury 2.5 hit general availability, GPT-6 Astra completed its commercial rollout, a major index re-ranked models, and OpenAI claimed a proof for a Millennium Prize problem. xAI slipped on Grok timing. All figures below are vendor- or third-party-reported, not independently verified.
The Week at a Glance
- Model launch: Mercury 2.5 (Inception) GA on Sep 8, diffusion-based LLM, 1,107 tok/s, 260K context, $0.20 / $0.75 per MTok.
- Full rollout: GPT-6 Astra generally available on ChatGPT tiers, API, Azure, and Bedrock. Enterprise plugins added.
- Benchmarks: Artificial Analysis Intelligence Index v4.3 published Sep 7. Top scores: Claude Fable 5.1 and GPT-6 Astra tie at 53.
- Science claims: OpenAI said an internal agent system solved the Navier-Stokes Millennium Prize problem (Sep 8). Google DeepMind released AlphaGenome Atlas.
- Pricing: Mercury 2.5 Preview discount expired; GA launch discount began.
Model Releases and Announcements
Mercury 2.5 GA (Inception Labs, Sep 8)
Inception launched the diffusion-based Mercury 2.5 as generally available on September 8. Claimed specs: 1,107 tok/s on NVIDIA GPUs, 260K context window, tunable reasoning. List price is $0.20 per million input tokens and $0.75 per million output tokens. Inception applied an 80% launch discount bringing it to $0.04 / $0.15.
Caveat: The Mercury 2.5 Preview discount expired September 8 at 07:00 UTC, reverting that endpoint to list price. Pricing tracker TokenCost notes Inception published no independent benchmarks for 2.5. (Sources: Inception blog, TokenCost, ORC Router.)
GPT-6 Astra Full Commercial Rollout (OpenAI)
GPT-6 Astra became generally available across ChatGPT Plus, Pro, Business, and Enterprise, plus the API (model name gpt-6-astra, $10 / $50 per MTok, ~$1 cached input, 2x for fast mode), Azure, and Bedrock. The rollout had started as a phased debut on September 3, but OpenAI's launch page is timestamped September 10. On September 9, Astra became the default for ChatGPT Work and Codex (admin opt-in). Enterprise plugins added include Oracle Analytics, Power BI, Navan, and Avalara. On September 11, OpenAI Developers advised reworking AGENTS.md and skills for Astra. (Sources: OpenAI launch page, Simon Carter, completeaitraining.com, OpenAI Developers.)
Benchmark Developments
Artificial Analysis Intelligence Index v4.3 (Sep 7)
The new index version ranks models on a composite score. Top-level results (max configuration with fallback where applicable):
| Model | Score | Cost per task |
|---|---|---|
| Claude Fable 5.1 (max w/ fallback) / GPT-6 Astra (max) | 53 (tie) | $7.63 / $3.26 |
| Claude Opus 5 | 51 | n/a |
| Claude Fable 5 | 50 | n/a |
| Muse Spark 1.3 | 48 | n/a |
| GPT-5.6 Sol | 47 | n/a |
Open-weight leaders: GLM 5.3 Flash (42), Qwen3.8 2.4T A95B (40), DeepSeek V4 Pro 0813 (36). Astra (max) costs $3.26 per task versus $7.63 for Fable 5.1 at equal score. (Source: AA article.)
Pricing, API Changes, and Provider News
- Mercury 2.5 pricing shift: Preview discount ended Sep 8; GA launch discount (80%) began. See above.
- Claude Code weekly limits: Scheduled to take effect September 14. Trackers describe the new limit as roughly 17% below current levels and 25% above the pre-May baseline. (Sources: aitoolsrecap, digitalapplied.)
- OpenAI Navier-Stokes claim (Sep 8): An internal agent system (~10,000 concurrent agents, 88 hours plus 17 hours of Lean verification) reportedly produced a proof of finite-time singularity for 3D Navier-Stokes, published as a written proof plus Lean formalization. OpenAI said the internal model was significantly more capable than GPT-6 Astra, will not claim the $1M Clay prize, and cannot rule out that de-identified product usage data aided training. NYU's Tristan Buckmaster and Anthropic's Levent Alpöge were working on a related Euler result; OpenAI's Sebastien Bubeck denied using their work. (Sources: Tao Media, Ground Truth, Guardian Mirror.)
- Google DeepMind AlphaGenome Atlas (Sep 8): Precomputed 1-petabyte catalogue of predicted effects of ~9 billion single-nucleotide variants; a new AVI score; free web portal for noncommercial use, plus API and Antigravity skill; commercial access via Google Cloud to follow; accompanied by a Nature paper. (Sources: Google blog, SiliconANGLE.)
What to Watch Next Week
- Claude Code weekly limits take effect September 14 (just outside this window, but relevant).
- Continued adoption of GPT-6 Astra across enterprise tiers.
- Any independent replication or commentary on the Navier-Stokes claim.
Takeaway: Mercury 2.5 gives the field a competitive diffusion-based option at a steep launch discount; GPT-6 Astra is now everywhere (and priced accordingly). The AA Index makes the cost-per-task gap between Fable 5.1 and Astra explicit: $7.63 vs $3.26 at the same score. The Navier-Stokes claim is heavy on promise, light on verifiable detail. Track the pricing, skip the hype.
For more context on earlier week's launches, see AI Model Roundup: Aug 17-23, 2026. For model rankings across categories, visit our Overall models page.