What the Top 10 Tells Us About Where AI Is Heading
I compared two lists this week: the top 10 models by actual usage on OpenRouter, and the top 10 by benchmark scores on Chatbot Arena.
Only six models appear on both lists. Four of the most-used models wouldnât crack a benchmark top 10. And two of the highest-scoring models barely make the usage top 30.
The gap between âbestâ and âmost usedâ is where the real story lives.
What the usage list has that the benchmark list doesnât
The usage top 10 skews toward:
- Cheaper models. Price per token correlates with usage volume more than any benchmark score does.
- Reliable models. Uptime and latency matter more than an extra 2% on MMLU.
- Ecosystem models. Models available on multiple platforms get used more â simple distribution math.
Developers optimize for predictability, not peak performance. The model that never goes down beats the model that occasionally writes a brilliant sonnet.
What the benchmark list has that the usage list doesnât
The benchmark top 10 skews toward:
- Cutting-edge releases. New models get benchmarked immediately. Adoption takes weeks or months.
- Niche specialists. Some models score incredibly well on specific tasks but lack broad appeal.
- Expensive flagships. The ââmoney is no objectâ tier. Great scores, low volume because the per-token cost is prohibitive.
Benchmarks measure potential. Usage measures practical value. Both matter. They just measure different things.
This weekâs ranks
Not much movement in the top 10 this week. A quiet week is its own signal â it means nobody launched anything disruptive and nobody had a major outage. In AI infrastructure, boring is beautiful.
The action is all in spots 15-30, where open-source models are jockeying for position. More on that next week.
How to use this
When someone tells you âModel X is the best,â ask: best by what measure?
If they mean benchmark scores, look at the usage rankings too. A model with great scores and low usage is sending you a signal â itâs either too expensive, too unreliable, or too hard to access.
On this site, we show both. Scores column. Price column. Platform availability. Make your own call.