Benchmarks × real traffic-weighted prices

The honest leaderboard
for LLM APIs.

The honest leaderboard for LLM APIs. Benchmarks from Artificial Analysis and llm-stats.com, measured against what people actually pay on OpenRouter. Models beaten on both price and quality get grayed out.

updated 2026-08-15 198 models tracked 11 worth buying 158 superseded 119 dropped
🪙
Budget
Ling-3.0-flash
30.4 pts · $0.023/M
💰
Best value
DeepSeek V4 Flash 0731
44.6 pts · $0.148/M
Best smart
DeepSeek V4 Pro 0813
48.2 pts · $0.478/M
🧠
Smartest
GPT-5.6 Sol (max)
54.7 pts · top quality
01 · The market

Price you actually pay vs. real quality

Every model on one plane — upper-left wins. Click a dot to open it on OpenRouter.

◆ Frontier (best at their price) (11)Near-frontier (close) (29)Grayed (superseded) (158)
02 · Value ladder

The best model at every price level

Every Pareto frontier point is a rung — anything not on a rung has a cheaper, better alternative.

03 · Lab leaderboard

Top labs by their strongest model

Premium = price of the lab's top model vs. the median lab's top model (×). 1× = going rate.

#LabBest modelPerfTop pricePremium◆ Mentions
1OpenAIGPT-5.6 Sol (max)54.7$27.39/M11.3×◆◆◆ 3
2AnthropicClaude Fable 553.8$11.61/M4.8×◆◆◆◆◆ 7
3Moonshot AIKimi K3 (max)52.6$8.58/M3.6× 1
4SpaceXAIGrok 4.650.3$2.41/M1.0×◆◆ 2
5QwenQwen3.8 Max50.0$2.22/M0.9×◆◆◆ 3
6DeepSeekDeepSeek V4 Pro 081348.2$0.478/M0.2×◆◆◆◆◆ 9
7GoogleGemini 3.7 Flash (high)47.9$0.615/M0.3×◆◆◆ 3
8MetaMuse Spark 1.146.8$1.36/M0.6×
03b · Lab momentum

Who's pulling ahead

Level = lab's top composite. Gained = cumulative perf raise across its releases this year.

◆ Pulling aheadleading & gaining
6
  • Qwentop 50.0 · +37.8
  • Moonshot AItop 52.6 · +30.7
  • OpenAItop 54.7 · +28.2
  • Anthropictop 53.8 · +25.1
  • DeepSeektop 48.2 · +29.3
  • Z.aitop 44.9 · +22.6
⬆️ Catching upgaining, not yet leading
3
  • Upstagetop 29.7 · +25.2
  • inclusionAItop 30.4 · +23.2
  • Mistral AItop 24.6 · +26.8
🎯 Leading, settledon top, flat
4
  • Googletop 47.9 · +16.8
  • SpaceXAItop 50.3 · +14.3
  • MiniMaxtop 38.8 · +15.4
  • Metatop 46.8 · +0.0
🛑 Stalledneither leading nor moving
9
  • NVIDIAtop 28.7 · +16.9
  • Tencenttop 38.2 · +5.1
  • Thinking Machinestop 34.8 · +1.2
  • Xiaomitop 33.3 · +0.0
  • KwaiPilottop 26.4 · +0.0
  • StepFuntop 21.1 · +0.5
  • Amazontop 14.3 · +2.9
  • Nous Researchtop 8.7 · +1.1
  • IBM Granitetop 0.8 · +0.0
04 · All models

Every model, sortable

Click headers to sort · grayed rows are models to skip · click a lab to filter.

Search
Lab