Inception: Mercury 2.5
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving... Parallel token generation reduces latency on short and medium responses Reasoning-oriented behavior at ultra-low /usr/bin/bash.04//usr/bin/bash.15 per million pricing 260K-token context supports long code and document tasks Diffusion decoding is a less mature serving approach than autoregressive models OpenRouter description and independent benchmark coverage are still limited Low-latency coding assistance High-volume reasoning workloads Long-context document and code analysis
Specifications
| Lab | Inception |
|---|---|
| Context window | 260,000 |
| Input price | $0.04/1M |
| Output price | $0.15/1M |
| Release | 2026-09-01 00:00:00 |
Strengths
- Parallel token generation reduces latency on short and medium responses
- Reasoning-oriented behavior at ultra-low /usr/bin/bash.04//usr/bin/bash.15 per million pricing
- 260K-token context supports long code and document tasks
Weaknesses
- Diffusion decoding is a less mature serving approach than autoregressive models
- OpenRouter description and independent benchmark coverage are still limited
Best for
- Low-latency coding assistance
- High-volume reasoning workloads
- Long-context document and code analysis
In Depth: Mercury 2.5
Summary
Mercury 2.5 is an AI model from Inception.
Released 2026-09-01 00:00:00. It supports Text input and produces Text output, with a context window of 260K tokens. Input pricing is $0.04 per 1M tokens and output is $0.15 per 1M tokens on OpenRouter.