Skip to content
LMRank
Models Categories Compare Tools Blog About
Search
Models Categories Compare Tools Blog About

LMRank Blog

AI model news, analysis, pricing context, and release notes.

AI Model Weekly Roundup - May 4–10, 2026
2026-05-11

AI Model Weekly Roundup - May 4–10, 2026

This week in AI: Mistral Medium 3.5 ships as a 128B open-weight flagship, Grok 4.3 stakes its 1M-context claim, and GPT-5.5 Instant undercuts the mid-tier.

Read more →
AI Model Weekly Roundup - May 18–24, 2026
2026-05-25

AI Model Weekly Roundup - May 18–24, 2026

This week in AI: DeepSeek's reasonix coding agent tops Hacker News, Google ships Gemini 3.5 Flash, xAI launches Grok Build 0.1, and memory costs climb.

Read more →
AI Model Weekly Roundup - May 25–31, 2026
2026-06-01

AI Model Weekly Roundup - May 25–31, 2026

This week in AI: DeepSWE crowns GPT-5.5 and flags Claude Opus 4.8, GitHub Copilot adds a 57x GPT-5.5 multiplier, and distillation drama hits Opus 4.8.

Read more →
Creative Benchmarks: Testing What AI Models Can Build
2026-05-09

Creative Benchmarks: Testing What AI Models Can Build

LMRank's Creative HTML Generation benchmarks test what AI models can actually build - landing pages, 3D simulations, and full projects from a single prompt.

Read more →
DeepSeek Makes Its 75% API Discount Permanent
2026-05-25

DeepSeek Makes Its 75% API Discount Permanent

DeepSeek confirmed its V4 Pro 75% API discount won't expire. At $0.50 per million tokens with a 9.0 LMRank score, this changes the economics of frontier AI.

Read more →
DeepSeek V4-Flash: Open-Weight, Million-Token Context
2026-05-19

DeepSeek V4-Flash: Open-Weight, Million-Token Context

DeepSeek V4-Flash brings a 1M-token context window to open-weight AI: 284B params, 13B active, MIT license, and coding scores that rival closed frontier models.

Read more →
GPT-5.5's Hidden Cost Surge: What OpenRouter Found
2026-05-10

GPT-5.5's Hidden Cost Surge: What OpenRouter Found

OpenRouter's analysis reveals GPT-5.5 costs 49-92% more in practice than earlier GPT-5 models. Here's the breakdown by prompt length and the best alternatives.

Read more →
Grok 4.3 vs Gemini Ultra 2: The Million-Token Showdown
2026-05-16

Grok 4.3 vs Gemini Ultra 2: The Million-Token Showdown

xAI's Grok 4.3 undercuts Gemini Ultra 2 by 90% on output pricing while matching its 1M token context window. Here's the breakdown.

Read more →
Kimi K2.6 vs Qwen3.6 35B: Open-Weight Coding Crown
2026-05-28

Kimi K2.6 vs Qwen3.6 35B: Open-Weight Coding Crown

Kimi K2.6 vs Qwen3.6 35B A3B: two open-weight coding models compared on price, LMRank score, and agentic reliability. The cheaper Qwen edges ahead on value.

Read more →
MiniMax M3: Open-Weight Coding, 1M Context, Multimodal
2026-06-01

MiniMax M3: Open-Weight Coding, 1M Context, Multimodal

MiniMax M3 ships open-weights with 1M context, native multimodal, and frontier coding claims at $0.30/M input tokens. Here's why it matters.

Read more →
Mistral Medium 3.5: 128B Open-Weight Agentic Coder
2026-05-31

Mistral Medium 3.5: 128B Open-Weight Agentic Coder

Mistral's new 128B dense model packs a 256K context window, open weights, and claims 77.6% on SWE-Bench. Here's why the medium tier just got interesting.

Read more →
Qwen3.7 Max: The $2.50 Sleeper Flagship
2026-05-22

Qwen3.7 Max: The $2.50 Sleeper Flagship

Qwen3.7 Max pairs a 9.0 LMRank score and 1M-token context with $2.50/M input pricing - one of the best-value frontier models you can call today.

Read more →
State of Large Language Models - May 2026
2026-05-15

State of Large Language Models - May 2026

The state of large language models in May 2026: who leads, who's overpriced, and which model to actually use - with live benchmark and pricing data from LMRank.

Read more →
When Claude Builds Claude: Inside Anthropic's Recursion
2026-06-07

When Claude Builds Claude: Inside Anthropic's Recursion

Anthropic says Claude now writes 80% of its own code and ships 8x faster - then called for a global AI pause. Inside the recursive self-improvement data.

Read more →
AI Model Weekly Roundup - June 2–8, 2026
2026-06-08

AI Model Weekly Roundup - June 2–8, 2026

This week in AI: Anthropic scales Claude Mythos into critical infrastructure, Google ships on-device Gemma 4 QAT, Meta delays again, and Linux wants Claude.

Read more →
Mixtral 8x22B v2: 512K Context, 40% Price Cut
2026-06-10

Mixtral 8x22B v2: 512K Context, 40% Price Cut

Mistral ships Mixtral 8x22B v2 with a 512K context window and 40% price cut, undercutting every long-context competitor on the market.

Read more →
The US Government Shut Down Claude Fable 5 in 3 Days
2026-06-13

The US Government Shut Down Claude Fable 5 in 3 Days

Anthropic launched Claude Fable 5 on June 9; a US export-control directive killed it worldwide by June 12 - the first government shutdown of a live AI API.

Read more →
The AI Price War of 2026: The Race to Zero
2026-06-13

The AI Price War of 2026: The Race to Zero

The 2026 AI price war is here: DeepSeek cut V4.1 15%, OpenAI added 90% cache discounts, and Google hit $0.10/M tokens. What falling prices mean for developers.

Read more →
AI Model Weekly Roundup - June 9–15, 2026
2026-06-15

AI Model Weekly Roundup - June 9–15, 2026

This week in AI: Claude Fable 5 was killed by a US export directive 72 hours after launch, Mistral shipped 512K-context Mixtral 8x22B v2, and prices fell.

Read more →
Kimi K2.7 Code: 1T Params, Self-Made Benchmarks
2026-06-16

Kimi K2.7 Code: 1T Params, Self-Made Benchmarks

Moonshot AI's Kimi K2.7 Code packs 1T parameters and a $0.95/M-token price, but every launch benchmark is proprietary. Here's why that matters.

Read more →
Model Routers: OpenRouter Fusion vs Sakana Fugu Ultra
2026-06-24

Model Routers: OpenRouter Fusion vs Sakana Fugu Ultra

Model routers are here: OpenRouter Fusion's parallel synthesis vs Sakana AI's learned Fugu Ultra orchestration - how they compare on latency, control, and cost.

Read more →
Google's 38% Gemini AI Plus Cut Fuels the Price War
2026-06-25

Google's 38% Gemini AI Plus Cut Fuels the Price War

Google slashed its AI Plus subscription from $7.99 to $4.99 per month while doubling storage. The consumer AI price war just got real.

Read more →
GPT-5.6 Sol/Terra/Luna: OpenAI's Gated Release
2026-06-28

GPT-5.6 Sol/Terra/Luna: OpenAI's Gated Release

OpenAI's GPT-5.6 Sol/Terra/Luna shipped only after US government preclearance - a tiered, gated release that could set the template for frontier AI launches.

Read more →
Weekly AI Model Roundup - June 22–29, 2026
2026-06-29

Weekly AI Model Roundup - June 22–29, 2026

This week in AI: OpenAI launches GPT-5.6 Sol under government gating, Claude Fable 5 stays offline, and Sakana Fugu Ultra hits frontier scores via routing.

Read more →
Model Routers Are the New Foundation: Comparing OpenRouter’s Optima and Sakana AI’s Fugu Ultra

Model Routers Are the New Foundation: Comparing OpenRouter’s Optima and Sakana AI’s Fugu Ultra

Read more →
When Your AI Gaslights You: Claude Opus 4.8's Hallucinated User Messages

When Your AI Gaslights You: Claude Opus 4.8's Hallucinated User Messages

Read more →
The AI Release Delay Controversy: Fable Blocked, GPT-5.6 Postponed

The AI Release Delay Controversy: Fable Blocked, GPT-5.6 Postponed

Read more →
Microsoft's Copilot Cowork Drops Per-Seat Pricing: The End of Flat-Rate Enterprise AI?

Microsoft's Copilot Cowork Drops Per-Seat Pricing: The End of Flat-Rate Enterprise AI?

Read more →
DeepSeek's No-Poaching Clause: The $50B AI Talent War Intensifies

DeepSeek's No-Poaching Clause: The $50B AI Talent War Intensifies

Read more →
LongCat-2.0: China's 1.6T-Parameter Coding Model Trained Without a Single Nvidia GPU
2026-07-01

LongCat-2.0: China's 1.6T-Parameter Coding Model Trained Without a Single Nvidia GPU

Meituan's LongCat-2.0 is a 1.6T-parameter MoE coding model trained entirely on domestic Chinese ASICs. It challenges the assumption that frontier AI requires Nvidia hardware.

Read more →
Ornith 1.0: The Open-Source Coding Model That Finally Beat Claude Opus 4.7

Ornith 1.0: The Open-Source Coding Model That Finally Beat Claude Opus 4.7

Read more →
Memora: Microsoft's 98% Token Reduction Could Unlock Truly Long-Running Agents

Memora: Microsoft's 98% Token Reduction Could Unlock Truly Long-Running Agents

Read more →
DeepSeek V4 Pro vs GPT-5.5 vs Claude Opus 4.8: Agentic Coding Price-to-Performance in 2026

DeepSeek V4 Pro vs GPT-5.5 vs Claude Opus 4.8: Agentic Coding Price-to-Performance in 2026

DeepSeek V4 Pro, GPT-5.5, and Claude Opus 4.8 compared head-to-head for agentic coding. Price, quality, and real engineering trade-offs for mid-2026 teams.

Read more →
Agentic Coding Showdown: DeepSeek V4 Pro vs GPT-5.5 vs Opus 4.8

Agentic Coding Showdown: DeepSeek V4 Pro vs GPT-5.5 vs Opus 4.8

DeepSeek V4 Pro, GPT-5.5, and Claude Opus 4.8 compared for agentic coding in mid-2026. Pricing, benchmarks, and which model to pick for real engineering work.

Read more →
AI Model Roundup: June 30–July 6, 2026
2026-07-06

AI Model Roundup: June 30–July 6, 2026

Claude Sonnet 5 launches with hidden cost increases. OpenAI previews GPT-5.6 Sol/Terra/Luna. Meituan open-sources LongCat 2.0. Plus new diffusion models and benchmark shifts.

Read more →
Tencent Hy3 vs DeepSeek V4 Pro: The New Open-Weight Price Fight
2026-07-06

Tencent Hy3 vs DeepSeek V4 Pro: The New Open-Weight Price Fight

Tencent Hy3 enters LMRank near DeepSeek V4 Pro with dramatically lower input pricing, strong agentic benchmarks, Apache 2.0 weights, and a clear role as a cheaper production worker model.

Read more →
DeepSeek Peak Pricing: The End of the AI Token Free Lunch
2026-07-07

DeepSeek Peak Pricing: The End of the AI Token Free Lunch

DeepSeek's shift from price cuts to peak-hour surge pricing reveals capacity constraints. What this means for inference economics and the 'race to zero.'

Read more →
When AI Agents Gaslight Themselves: Claude Code Confabulates a Fake Security Attack

When AI Agents Gaslight Themselves: Claude Code Confabulates a Fake Security Attack

Read more →
China Bans Emotional AI: The West's Wild West vs. the World's First Emotional Interaction Law

China Bans Emotional AI: The West's Wild West vs. the World's First Emotional Interaction Law

Read more →
Metered Billing Ends the Flat-Rate AI Era
2026-07-10

Metered Billing Ends the Flat-Rate AI Era

Anthropic, OpenAI, and Meta have all shifted flagship models and agents to metered billing. Users must adopt token-cost discipline or pay more.

Read more →
Meta’s $4.25 Output Price: The New Bottom for Frontier API Pricing

Meta’s $4.25 Output Price: The New Bottom for Frontier API Pricing

Read more →
GhostApproval & GitLost: Why Agentic Coding Is Breakable by Design

GhostApproval & GitLost: Why Agentic Coding Is Breakable by Design

Read more →
GPT-5.6 Sol, Terra, Luna: Compared, and the One Rule
2026-07-10

GPT-5.6 Sol, Terra, Luna: Compared, and the One Rule

OpenAI shipped GPT-5.6 as three tiers with a reasoning-effort dial. The benchmark data says it is really three models at max effort. Here is the comparison and the rule of thumb.

Read more →
The High Cost of Fluency: Why General AI Models Fail Medical Grounding

The High Cost of Fluency: Why General AI Models Fail Medical Grounding

Read more →
xAI’s Government Pivot: Grok 4.5 as a Defense Contract Bid

xAI’s Government Pivot: Grok 4.5 as a Defense Contract Bid

Read more →
OpenAI’s Cerebras Lock-In: The End of Hardware Agnostic API Speed

OpenAI’s Cerebras Lock-In: The End of Hardware Agnostic API Speed

Read more →
Alibaba Ditches Hybrid Reasoning: Qwen3’s New Efficiency Play

Alibaba Ditches Hybrid Reasoning: Qwen3’s New Efficiency Play

Read more →
Thinking Machines Inkling: Open Weights, Not Just Another Leaderboard Model
2026-07-18

Thinking Machines Inkling: Open Weights, Not Just Another Leaderboard Model

Thinking Machines Lab's first model, Inkling, trades top-line leaderboard performance for something harder to buy: an open multimodal base that developers can fine-tune and reshape.

Read more →
Kimi K3 Pricing Ends the 'Open Weights = Cheap' Era
2026-07-19

Kimi K3 Pricing Ends the 'Open Weights = Cheap' Era

Moonshot AI's Kimi K3 launch at $3/$15 per MTok — matching Anthropic's Sonnet tier — shatters the assumption that open-weight models are budget options.

Read more →
Grok 4.5’s $2/$6 Price Tag vs. Databricks’ Gateway Pivot

Grok 4.5’s $2/$6 Price Tag vs. Databricks’ Gateway Pivot

Read more →
The Safety Tax of Agentic AI: OpenAI’s GPT-5.6 Sol admits to 'Honest Mistakes'

The Safety Tax of Agentic AI: OpenAI’s GPT-5.6 Sol admits to 'Honest Mistakes'

Read more →
AI Model Roundup: July 13-19, 2026

AI Model Roundup: July 13-19, 2026

Read more →
Kimi K3 and China's AI Offensive: Weekly Roundup July 13-19
2026-07-20

Kimi K3 and China's AI Offensive: Weekly Roundup July 13-19

Moonshot AI's Kimi K3 launch, Alibaba's Qwen3.8 Max Preview, and Thinking Machines Inkling signal a decisive shift in open-weight AI leadership during WAIC 2026.

Read more →
The Sandbox Paradox: Why Hugging Face’s Breach Proved Open Weights Are Critical for Defense

The Sandbox Paradox: Why Hugging Face’s Breach Proved Open Weights Are Critical for Defense

Read more →
The New Price War: How Google’s Token Efficiency Undercuts Raw Performance

The New Price War: How Google’s Token Efficiency Undercuts Raw Performance

Read more →
The Long-Horizon Trap: Why Cheap Open Models Beat $10K GPT Instances on Agent Workloads

The Long-Horizon Trap: Why Cheap Open Models Beat $10K GPT Instances on Agent Workloads

Read more →
AI Model Roundup: Aug 17-23, 2026
2026-08-28

AI Model Roundup: Aug 17-23, 2026

Open-weight releases from Z.AI, Ornith AI, and Alibaba compete on coding benchmarks this week. Pricing shifts from OpenAI and DeepSeek change the cost landscape.

Read more →
The Hallucination Trade-Off: Agent Reliability Replaces Benchmark Chasing in 2026
2026-08-28

The Hallucination Trade-Off: Agent Reliability Replaces Benchmark Chasing in 2026

Chinese labs decouple quality from hallucination rates as Grok 4.6, Qwen3.8-Flash-Next, and GLM-5.3-Flash shift the war from cheapest model to most reliable agent infrastructure.

Read more →
From All-You-Can-Eat to Load Balancing: The End of Flat-Rate AI Subscriptions

From All-You-Can-Eat to Load Balancing: The End of Flat-Rate AI Subscriptions

Read more →
Benchmark Saturation: Why Anthropic Is Killing GPQA Diamond and What Replaces It

Benchmark Saturation: Why Anthropic Is Killing GPQA Diamond and What Replaces It

Read more →

Get new rankings and analysis by email

LMRank

Independent AI model rankings. Compare LLMs by benchmark score, pricing, and real-world fit.

Product

Leaderboard Model finder Categories Tools Agents Compare models Showcases

Resources

Blog Methodology About Sitemap

© 2026 LMRank. All rights reserved.