Benchmarks · 2026-09-12
Do the best-scoring models win?
Independent LLM benchmark scores from OpenRouter's unified endpoint (Artificial Analysis composite indexes), cross-referenced with the token share those AI models receive on OpenRouter today.
Top 20 models — Intelligence
Higher is better. Bars coloured by lab.
Efficiency frontier — Intelligence vs output price
Bubble = usage share. The line traces the Pareto frontier: models where nothing cheaper scores as high.
How to read: X-axis is list output price (log); Y-axis is benchmark score; bubble area is current token share. The connecting line joins models that dominate on price-for-score.
Benchmark score vs OpenRouter share — Intelligence
Top-right = high score AND heavily used · bottom-right = high score but under-adopted · top-left = popular despite lower score. Bubble size scales with usage share.
What a session actually costs in Hermes Agent
Median cost of one coding session by session length. 30-day window, log scale.
Claude Fable 5 costs 8614× more per session than Mistral Nemo in this agent.
How to read: Each row is a model. Dots run from short sessions to long ones — the further right, the more a median session costs. The number on the right is the 10–49 turns median.
- $4.22
- $4.11
- $3.98
- $3.40
- $2.18
- $1.75
- $1.65
- $1.24
- $1.18
- $0.934
- $0.785
- $0.785
Intelligence per dollar of real work
Benchmark score against the median 10–49 turns session cost in Hermes Agent.
Models on the dashed line are undominated: nothing is both cheaper and better-scoring. Bubble size is OpenRouter token share.
How to read: Cost runs left to right on a log scale, benchmark score bottom to top. Up and to the left is better value. Click a bubble for the model page.