How this comparison works
A daily-rebuilt comparison of frontier AI models. Every number is sourced from public benchmarks, and no large language model writes any of this content.
Where the data comes from
All benchmark scores, pricing, context windows, and speed figures are pulled from the public llm-stats.com leaderboard, which aggregates results from verified public benchmarks and live API metrics. The comparison refreshes once a day, so the numbers you are reading reflect the latest published results.
The leaderboard tracks over 300 models and 50 benchmarks. We select a focused matrix of the ten most relevant frontier models for builders and operators. The selection is curated to include the top reasoning models, the strongest open-weight contenders, and the models a production team would actually evaluate side by side.
How scores are computed
Each benchmark reports a raw score and a maximum. We normalize every result to a 0 to 1 ratio (score divided by maximum), then group benchmarks by workload category: reasoning, coding, agents and tool use, math, vision and multimodal, long context, writing, and knowledge. A model's index on each axis is the average of its normalized scores across the benchmarks tagged with that category. When a model has not published a result for a benchmark, that benchmark is shown as "not reported" and does not drag the average down.
This means a model with fewer published results is not penalized, but it also means you should treat an index backed by only one or two benchmarks with more caution than one backed by a dozen. The benchmark count on each model page tells you how many results went into its scores.
Pricing and speed
The blended price is the input and output cost per million tokens combined at an 8:1 input-to-output mix, matching the convention used by the source leaderboard. This ratio reflects a typical chat or agent workload where the prompt is longer than the response. Speed figures come from the provider's measured throughput and latency where available, supplemented by the leaderboard's live metrics feed which samples actual API response times.
How the verdicts are generated
The plain-English verdict on each comparison page is computed with simple arithmetic from the numbers above: which model is cheaper, and which axis shows the largest normalized gap. The "which should you choose" guidance applies the same rules to a fixed set of common workloads: hard reasoning, coding, agents, math, vision, long documents, writing, cost-sensitive high-volume work, and low-latency interactive use.
Nothing here is generated by an LLM. The wording is always reproducible and traceable back to the underlying numbers. If the data changes, the verdict changes with it, deterministically.
Honesty and limitations
- Benchmarks are imperfect. A single number never captures a model's real usefulness on your specific task, with your specific data, at your specific scale. Always test on your own workload before committing.
- Some results are self-reported. The source leaderboard labels self-reported scores. We include them because they are often the only published result for a new model, but verified scores carry more weight.
- Prices and speeds change. This page is refreshed daily, but always confirm against the provider's current pricing page before committing budget. API latency varies by region, time of day, and provider load.
- Benchmark contamination is real. Public benchmarks leak into training data. A high score on a widely published test does not guarantee the model generalizes to novel problems in the same domain.