The honest leaderboard for LLM APIs. Benchmarks from Artificial Analysis and llm-stats.com, measured against what people actually pay on OpenRouter. Models beaten on both price and quality get grayed out.
Every model on one plane — upper-left wins. Click a dot to open it on OpenRouter.
Every Pareto frontier point is a rung — anything not on a rung has a cheaper, better alternative.
Premium = price of the lab's top model vs. the median lab's top model (×). 1× = going rate.
| # | Lab | Best model | Perf | Top price | Premium | ◆ Mentions |
|---|---|---|---|---|---|---|
| 1 | 54.7 | $27.39/M | 11.3× | ◆◆◆ 3 | ||
| 2 | 53.8 | $11.61/M | 4.8× | ◆◆◆◆◆ 7 | ||
| 3 | 52.6 | $8.58/M | 3.6× | ◆ 1 | ||
| 4 | 50.3 | $2.41/M | 1.0× | ◆◆ 2 | ||
| 5 | 50.0 | $2.22/M | 0.9× | ◆◆◆ 3 | ||
| 6 | 48.2 | $0.478/M | 0.2× | ◆◆◆◆◆ 9 | ||
| 7 | 47.9 | $0.615/M | 0.3× | ◆◆◆ 3 | ||
| 8 | 46.8 | $1.36/M | 0.6× | — |
Level = lab's top composite. Gained = cumulative perf raise across its releases this year.
Click headers to sort · grayed rows are models to skip · click a lab to filter.