by datastudy.nl

Head-to-head comparisons of frontier AI models, rebuilt daily from public benchmark data

Data Today comparison

LLM Compare

How the top frontier models stack up: benchmark scores, pricing, speed, and which one to pick for your workload. Every number is sourced from public benchmarks and refreshed daily.


Why this exists

Every week a new frontier model ships, and every launch comes with a different set of cherry-picked benchmarks. One provider leads with GPQA, another with SWE-Bench, a third with their own internal eval. This comparison pulls every public result into one place, normalizes them to the same scale, and tells you which model actually wins on the workload you care about. No marketing, no vibes, just the numbers.

The data comes from the llm-stats.com public leaderboard, which aggregates results from verified benchmarks and live API metrics. We refresh it daily. The verdicts and recommendations are computed with plain arithmetic from those numbers: nothing here is written by an LLM.

Pick two models to compare

Select two models


Recent comparisons


From the Data Today newsroom

Data and AI, told like a newsroom.

This comparison is part of Data Today. Read the wider data and AI coverage on the main site.

Go to Data Today