Inception: Mercury 2.5 Preview
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving... Parallel token generation targets very low latency Reasoning capability at ultra-low token pricing 260K-token context window Useful for fast coding and structured generation Preview release may change behavior or availability Diffusion generation is less established than autoregressive models Latency-sensitive coding assistants High-volume structured generation Fast reasoning prototypes
Specifications
| Lab | Inception |
|---|---|
| Context window | 260,000 |
| Input price | $0.04/1M |
| Output price | $0.15/1M |
| Release | 2026-08-01 00:00:00 |
Strengths
- Parallel token generation targets very low latency
- Reasoning capability at ultra-low token pricing
- 260K-token context window
- Useful for fast coding and structured generation
Weaknesses
- Preview release may change behavior or availability
- Diffusion generation is less established than autoregressive models
Best for
- Latency-sensitive coding assistants
- High-volume structured generation
- Fast reasoning prototypes
In Depth: Mercury 2.5 Preview
Summary
Mercury 2.5 Preview is an AI model from Inception.
Released 2026-08-01 00:00:00. It supports Text input and produces Text output, with a context window of 260K tokens. Input pricing is $0.04 per 1M tokens and output is $0.15 per 1M tokens on OpenRouter.