Inclusionai: Ling-3.0-flash
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
Specifications
| Attribute | Value |
|---|---|
| Lab | Inclusionai |
| Tags | Intelligent Agentic Fast Open Weight |
| Release Date | 2026-07 |
| Context Window | 262,144 tokens |
| Input Price / 1M | $0.02 |
| Output Price / 1M | $0.06 |
| Input Modalities | Text |
| Output Modalities | Text |
Strengths
- 256K-token context window for long documents
- Low input pricing for cost-sensitive production workloads
- Tool-use support for agentic workflows
Weaknesses
- Newer listing with limited independent benchmark coverage
- Text-only inputs limit multimodal application coverage
Best For
- Agentic workflows and tool-using applications
- Long-context document and codebase processing
In Depth: Ling-3.0-flash
Summary
Ling-3.0-flash is an AI model from Inclusionai.
Released 2026-07. It supports Text input and produces Text output, with a context window of 262,144 tokens. Input pricing is $0.02 per 1M tokens and output is $0.06 per 1M tokens on OpenRouter.