by datastudy.nl

Head-to-head comparisons of frontier AI models, rebuilt daily from public benchmark data

The models

All models compared

Every model in the current comparison matrix. Tap a model to see its full profile, workload strengths, and every head-to-head matchup against the rest of the field.

The ten models below represent the frontier of large language model capability as of today. They span four countries, three of them are open-weight, and their blended prices range from under a dollar to nearly eight dollars per million tokens. The comparison matrix covers every ordered pair: 90 head-to-head pages in total, each with per-benchmark scores, workload recommendations, and a plain-English verdict computed from the numbers.

The selection is curated to keep the matrix focused on models a builder would actually consider for production work. It includes the top reasoning models from OpenAI, Anthropic, and Moonshot AI, the leading open-weight contender from DeepSeek, the strongest Chinese closed models from Alibaba and Zhipu, and xAI's Grok. The full list refreshes when the underlying leaderboard data changes.


Each model's benchmark scores are normalized to a 0 to 1 scale from public results, and averaged per workload axis. Pricing is the blended input/output cost per million tokens at an 8:1 mix. How this works.