Browse by Category
Find the best model for the job you actually have. Each category is a curated, independently ranked shortlist - top picks with live pricing and rank, not a single all-purpose winner.
Pick a category
Full model listOverall
113 modelsTop-ranked models across our benchmark mix.
- Claude Fable 5
- Claude Opus 4.8
- Claude Opus 4.7
Coding
3 modelsTop picks for code generation, debugging, and refactoring.
- Claude Fable 5
- DeepSeek V4 Pro preview
- GPT-5.3-Codex
Agentic Coding Model
3 modelsTop picks for agentic, multi-step coding workflows.
- Claude Sonnet 4.6
- DeepSeek V4 Pro preview
- GLM 5.2
Open-Weight
3 modelsOpen-weight models you can self-host or inspect.
- DeepSeek V4 Pro preview
- Z.ai: GLM 5
- Qwen3 235B A22B
Cheap Model
3 modelsStrong models with the lowest per-token pricing.
- DeepSeek V4 Pro preview
- Hy3 preview
- gpt-oss-120b
Multimodal
3 modelsModels that handle image, audio, and text inputs.
- Gemini 3.1 Pro Preview
- Z.ai: GLM 5V Turbo
- MiMo-V2.5
Long Context
3 modelsModels with the largest usable context windows.
- Grok 4.20
- GLM 5.2
- Llama 4 Scout
Reasoning
3 modelsModels tuned for step-by-step reasoning and math.
- DeepSeek R1
- Grok 4.3
- o4 Mini
Fast Model
3 modelsTurbo, Flash, Mini, and Instant variants tuned for latency.
- Gemini 3.5 Flash
- Claude Haiku 4.5
- GPT-5.4 Mini
Local Model
3 modelsSmallest open-weight runs that fit on a single machine.
- LFM2.5-8B-A1B
- Qwen3.5-9B
- Gemma 4 26B A4B
Keep exploring
AI model category FAQs
How are the categories on LMRank ranked?
How often do category rankings update?
Can one model appear in more than one category?
Which categories does LMRank track?
About the LMRank AI model categories
Every model on LMRank is classified into a focused category that reflects its primary strength. Categories span reasoning, code generation, agentic tool use, retrieval, and long-context work. They are not assigned by the models' creators or their marketing copy. Instead, each category is built from structured, reproducible benchmarks that test models on real tasks within that domain, then ranks them by how well they perform against their peers.
Category rankings refresh automatically as new models ship and our benchmark suite expands. Because every category is tied to live evaluation data, the leaderboard in each category reflects the latest measurable performance, not a static editorial ranking. When a new model breaks into the top tier, the category page updates without manual curation. This keeps the comparisons actionable for teams evaluating models against production workloads rather than general-purpose leaderboards.
Each category page surfaces the top-ranked models with their pricing, provider, and current rank so you can reason about cost and capability together. Models can appear in multiple categories when they perform well across domains, but every category is ranked independently. A model that tops the coding leaderboard may sit mid-pack in creative writing, and that's intentional. The categories exist to help you find the right tool for the job, not to crown a single winner.