Category
Fast Model Models
The lowest-latency variants from each provider — Flash, Turbo, Mini, and Instant models — for interactive apps.
3 models
← All categories
Top Fast Model models
| Model | Pricing |
|---|---|
| Gemini 3.5 Flash Google DeepMind | $1.50/1M input |
| Claude Haiku 4.5 Anthropic | $1.00/1M input |
| GPT-5.4 Mini OpenAI | $0.75/1M input |
How these models were selected
These are smaller or latency-oriented model tiers intended for interactive use. Placement considers model family, pricing, and available public speed evidence.
When to use a different shortlist
Fast variants usually trade away difficult reasoning and long-horizon tool reliability. Measure end-to-end task completion, not tokens per second alone.
Open each model page to verify current pricing, context limits, source links, and known limitations before choosing a provider.