LLM Compare
How the top frontier models stack up: benchmark scores, pricing, speed, and which one to pick for your workload. Every number is sourced from public benchmarks and refreshed daily.
Why this exists
Every week a new frontier model ships, and every launch comes with a different set of cherry-picked benchmarks. One provider leads with GPQA, another with SWE-Bench, a third with their own internal eval. This comparison pulls every public result into one place, normalizes them to the same scale, and tells you which model actually wins on the workload you care about. No marketing, no vibes, just the numbers.
The data comes from the llm-stats.com public leaderboard, which aggregates results from verified benchmarks and live API metrics. We refresh it daily. The verdicts and recommendations are computed with plain arithmetic from those numbers: nothing here is written by an LLM.
Pick two models to compare
Select two models
Recent comparisons
GPT-5.6 Sol vs Claude Opus 5
GPT-5.6 Sol leads on 10 of 12 metrics
GPT-5.6 Sol vs Kimi K3
GPT-5.6 Sol leads on 7 of 13 metrics
GPT-5.6 Sol vs Claude Fable 5
GPT-5.6 Sol leads on 11 of 13 metrics
GPT-5.6 Sol vs GPT-5.6 Terra
GPT-5.6 Sol leads on 12 of 12 metrics
GPT-5.6 Sol vs Qwen3.8 Max
GPT-5.6 Sol leads on 7 of 13 metrics
GPT-5.6 Sol vs Grok 4.5
GPT-5.6 Sol leads on 12 of 12 metrics
GPT-5.6 Sol vs DeepSeek-V4-Flash-0731
GPT-5.6 Sol leads on 11 of 12 metrics
GPT-5.6 Sol vs GLM-5.2
GPT-5.6 Sol leads on 10 of 12 metrics
GPT-5.6 Sol vs Claude Sonnet 5
GPT-5.6 Sol leads on 11 of 14 metrics
Claude Opus 5 vs GPT-5.6 Sol
GPT-5.6 Sol leads on 10 of 12 metrics
Claude Opus 5 vs Kimi K3
Kimi K3 leads on 10 of 11 metrics
Claude Opus 5 vs Claude Fable 5
Claude Fable 5 leads on 5 of 9 metrics
Claude Opus 5 vs GPT-5.6 Terra
GPT-5.6 Terra leads on 9 of 12 metrics
Claude Opus 5 vs Qwen3.8 Max
Qwen3.8 Max leads on 9 of 12 metrics
Claude Opus 5 vs Grok 4.5
Grok 4.5 leads on 5 of 9 metrics
Claude Opus 5 vs DeepSeek-V4-Flash-0731
Claude Opus 5 leads on 6 of 8 metrics
Claude Opus 5 vs GLM-5.2
GLM-5.2 leads on 5 of 9 metrics
Claude Opus 5 vs Claude Sonnet 5
Claude Sonnet 5 leads on 7 of 10 metrics
From the Data Today newsroom
Data and AI, told like a newsroom.
This comparison is part of Data Today. Read the wider data and AI coverage on the main site.
Go to Data Today