Inclusionai: Ling-2.6-flash
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency.... 7.4B active params for extremely fast inference Ultra-low cost at $0.01/$0.03 per 1M tokens Strong token efficiency for coding and docs Performance competitive with same-scale state-of-the-art models Limited reasoning depth vs full-scale models 262K context window, smaller than 1T variant Not suited for complex multi-step agent tasks Real-time agent workflows Cost-sensitive coding and code review Document processing and summarization Lightweight automation pipelines
Specifications
| Lab | Inclusionai |
|---|---|
| Context window | 262,144 |
| Input price | $0.01/1M |
| Output price | $0.03/1M |
| Release | 2026-04-01 00:00:00 |
Strengths
- 7.4B active params for extremely fast inference
- Ultra-low cost at $0.01/$0.03 per 1M tokens
- Strong token efficiency for coding and docs
- Performance competitive with same-scale state-of-the-art models
Weaknesses
- Limited reasoning depth vs full-scale models
- 262K context window, smaller than 1T variant
- Not suited for complex multi-step agent tasks
Best for
- Real-time agent workflows
- Cost-sensitive coding and code review
- Document processing and summarization
- Lightweight automation pipelines
In Depth: Ling-2.6-flash
Summary
Ling-2.6-flash is an AI model from Inclusionai.
Released 2026-04-01 00:00:00. It currently appears in the Overall category on LMRank. It supports Text input and produces Text output, with a context window of 262.1K tokens. Input pricing is $0.01 per 1M tokens and output is $0.03 per 1M tokens on OpenRouter.