
AI Model Weekly Roundup - May 4–10, 2026
This week in AI: Mistral Medium 3.5 ships as a 128B open-weight flagship, Grok 4.3 stakes its 1M-context claim, and GPT-5.5 Instant undercuts the mid-tier.
Read more →AI model news, analysis, pricing context, and release notes.

This week in AI: Mistral Medium 3.5 ships as a 128B open-weight flagship, Grok 4.3 stakes its 1M-context claim, and GPT-5.5 Instant undercuts the mid-tier.
Read more →
This week in AI: DeepSeek's reasonix coding agent tops Hacker News, Google ships Gemini 3.5 Flash, xAI launches Grok Build 0.1, and memory costs climb.
Read more →
This week in AI: DeepSWE crowns GPT-5.5 and flags Claude Opus 4.8, GitHub Copilot adds a 57x GPT-5.5 multiplier, and distillation drama hits Opus 4.8.
Read more →
LMRank's Creative HTML Generation benchmarks test what AI models can actually build - landing pages, 3D simulations, and full projects from a single prompt.
Read more →
DeepSeek confirmed its V4 Pro 75% API discount won't expire. At $0.50 per million tokens with a 9.0 LMRank score, this changes the economics of frontier AI.
Read more →
DeepSeek V4-Flash brings a 1M-token context window to open-weight AI: 284B params, 13B active, MIT license, and coding scores that rival closed frontier models.
Read more →
OpenRouter's analysis reveals GPT-5.5 costs 49-92% more in practice than earlier GPT-5 models. Here's the breakdown by prompt length and the best alternatives.
Read more →
xAI's Grok 4.3 undercuts Gemini Ultra 2 by 90% on output pricing while matching its 1M token context window. Here's the breakdown.
Read more →
Kimi K2.6 vs Qwen3.6 35B A3B: two open-weight coding models compared on price, LMRank score, and agentic reliability. The cheaper Qwen edges ahead on value.
Read more →
MiniMax M3 ships open-weights with 1M context, native multimodal, and frontier coding claims at $0.30/M input tokens. Here's why it matters.
Read more →
Mistral's new 128B dense model packs a 256K context window, open weights, and claims 77.6% on SWE-Bench. Here's why the medium tier just got interesting.
Read more →
Qwen3.7 Max pairs a 9.0 LMRank score and 1M-token context with $2.50/M input pricing - one of the best-value frontier models you can call today.
Read more →
The state of large language models in May 2026: who leads, who's overpriced, and which model to actually use - with live benchmark and pricing data from LMRank.
Read more →
Anthropic says Claude now writes 80% of its own code and ships 8x faster - then called for a global AI pause. Inside the recursive self-improvement data.
Read more →
This week in AI: Anthropic scales Claude Mythos into critical infrastructure, Google ships on-device Gemma 4 QAT, Meta delays again, and Linux wants Claude.
Read more →
Mistral ships Mixtral 8x22B v2 with a 512K context window and 40% price cut, undercutting every long-context competitor on the market.
Read more →
Anthropic launched Claude Fable 5 on June 9; a US export-control directive killed it worldwide by June 12 - the first government shutdown of a live AI API.
Read more →
The 2026 AI price war is here: DeepSeek cut V4.1 15%, OpenAI added 90% cache discounts, and Google hit $0.10/M tokens. What falling prices mean for developers.
Read more →
This week in AI: Claude Fable 5 was killed by a US export directive 72 hours after launch, Mistral shipped 512K-context Mixtral 8x22B v2, and prices fell.
Read more →
Moonshot AI's Kimi K2.7 Code packs 1T parameters and a $0.95/M-token price, but every launch benchmark is proprietary. Here's why that matters.
Read more →
Model routers are here: OpenRouter Fusion's parallel synthesis vs Sakana AI's learned Fugu Ultra orchestration - how they compare on latency, control, and cost.
Read more →
Google slashed its AI Plus subscription from $7.99 to $4.99 per month while doubling storage. The consumer AI price war just got real.
Read more →
OpenAI's GPT-5.6 Sol/Terra/Luna shipped only after US government preclearance - a tiered, gated release that could set the template for frontier AI launches.
Read more →
This week in AI: OpenAI launches GPT-5.6 Sol under government gating, Claude Fable 5 stays offline, and Sakana Fugu Ultra hits frontier scores via routing.
Read more →





Meituan's LongCat-2.0 is a 1.6T-parameter MoE coding model trained entirely on domestic Chinese ASICs. It challenges the assumption that frontier AI requires Nvidia hardware.
Read more →


DeepSeek V4 Pro, GPT-5.5, and Claude Opus 4.8 compared head-to-head for agentic coding. Price, quality, and real engineering trade-offs for mid-2026 teams.
Read more →
DeepSeek V4 Pro, GPT-5.5, and Claude Opus 4.8 compared for agentic coding in mid-2026. Pricing, benchmarks, and which model to pick for real engineering work.
Read more →
Claude Sonnet 5 launches with hidden cost increases. OpenAI previews GPT-5.6 Sol/Terra/Luna. Meituan open-sources LongCat 2.0. Plus new diffusion models and benchmark shifts.
Read more →
Tencent Hy3 enters LMRank near DeepSeek V4 Pro with dramatically lower input pricing, strong agentic benchmarks, Apache 2.0 weights, and a clear role as a cheaper production worker model.
Read more →
DeepSeek's shift from price cuts to peak-hour surge pricing reveals capacity constraints. What this means for inference economics and the 'race to zero.'
Read more →


Anthropic, OpenAI, and Meta have all shifted flagship models and agents to metered billing. Users must adopt token-cost discipline or pay more.
Read more →


OpenAI shipped GPT-5.6 as three tiers with a reasoning-effort dial. The benchmark data says it is really three models at max effort. Here is the comparison and the rule of thumb.
Read more →




Thinking Machines Lab's first model, Inkling, trades top-line leaderboard performance for something harder to buy: an open multimodal base that developers can fine-tune and reshape.
Read more →
Moonshot AI's Kimi K3 launch at $3/$15 per MTok — matching Anthropic's Sonnet tier — shatters the assumption that open-weight models are budget options.
Read more →



Moonshot AI's Kimi K3 launch, Alibaba's Qwen3.8 Max Preview, and Thinking Machines Inkling signal a decisive shift in open-weight AI leadership during WAIC 2026.
Read more →



Open-weight releases from Z.AI, Ornith AI, and Alibaba compete on coding benchmarks this week. Pricing shifts from OpenAI and DeepSeek change the cost landscape.
Read more →
Chinese labs decouple quality from hallucination rates as Grok 4.6, Qwen3.8-Flash-Next, and GLM-5.3-Flash shift the war from cheapest model to most reliable agent infrastructure.
Read more →
