Mercury 2.5
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception.
Specifications
| Lab | Inception |
|---|---|
| Context window | 260K tokens |
| Input price | $0.04/1M |
| Output price | $0.15/1M |
| Release | Sep 2026 |
Strengths
- Parallel token generation reduces latency on short and medium responses
- Reasoning-oriented behavior at ultra-low /usr/bin/bash.04//usr/bin/bash.15 per million pricing
- 260K-token context supports long code and document tasks
Weaknesses
- Diffusion decoding is a less mature serving approach than autoregressive models
- OpenRouter description and independent benchmark coverage are still limited
Best for
- Low-latency coding assistance
- High-volume reasoning workloads
- Long-context document and code analysis
In Depth: Mercury 2.5
Summary
Mercury 2.5 is an AI model from Inception.
Released in Sep 2026. It takes text input and produces text output, with a context window of 260K tokens. Input pricing is $0.04 per 1M tokens and output is $0.15 per 1M tokens on OpenRouter.