All models compared
Every model in the current comparison matrix. Tap a model to see its full profile, workload strengths, and every head-to-head matchup against the rest of the field.
The ten models below represent the frontier of large language model capability as of today. They span four countries, three of them are open-weight, and their blended prices range from under a dollar to nearly eight dollars per million tokens. The comparison matrix covers every ordered pair: 90 head-to-head pages in total, each with per-benchmark scores, workload recommendations, and a plain-English verdict computed from the numbers.
The selection is curated to keep the matrix focused on models a builder would actually consider for production work. It includes the top reasoning models from OpenAI, Anthropic, and Moonshot AI, the leading open-weight contender from DeepSeek, the strongest Chinese closed models from Alibaba and Zhipu, and xAI's Grok. The full list refreshes when the underlying leaderboard data changes.
GPT-5.6 Sol
Proprietary · Multimodal · 1.05M ctx · $7.78/1M
Claude Opus 5
Proprietary · Multimodal · 1M ctx · $7.22/1M
Kimi K3
Open weights · Multimodal · 1.05M ctx · $4.33/1M
Claude Fable 5
Proprietary · Multimodal · 1M ctx · $14.44/1M
GPT-5.6 Terra
Proprietary · Multimodal · 1.05M ctx · $3.11/1M
Qwen3.8 Max
Proprietary · Multimodal · n/a ctx
Grok 4.5
Proprietary · Multimodal · 500K ctx · $2.44/1M
DeepSeek-V4-Flash-0731
Open weights · 1.05M ctx · $0.10/1M
GLM-5.2
Open weights · 1.05M ctx · $1.18/1M
Claude Sonnet 5
Proprietary · Multimodal · 1M ctx · $2.89/1M
Each model's benchmark scores are normalized to a 0 to 1 scale from public results, and averaged per workload axis. Pricing is the blended input/output cost per million tokens at an 8:1 mix. How this works.