Category
Multimodal Models
The strongest multimodal models — those that can take image, audio, or video in alongside text.
3 models
← All categories
Top Multimodal models
| Model | Pricing |
|---|---|
| Gemini 3.1 Pro Preview Google DeepMind | $2.00/1M input |
| Z.ai: GLM 5V Turbo Z Ai | $1.20/1M input |
| MiMo-V2.5 Xiaomi | $0.14/1M input |
How these models were selected
Models must accept at least one non-text input modality and provide enough general capability to analyze mixed text and media in practical workflows.
When to use a different shortlist
Modality support varies by provider and endpoint, and audio or video limits can differ from text limits. Verify the exact API before designing around a feature.
Open each model page to verify current pricing, context limits, source links, and known limitations before choosing a provider.