Benchmarks · 2026-09-02
Do the best-scoring models win?
Independent LLM benchmark scores from OpenRouter's unified endpoint (Artificial Analysis composite indexes), cross-referenced with the token share those AI models receive on OpenRouter today.
Top 20 models — Intelligence
Higher is better. Bars coloured by lab.
Efficiency frontier — Intelligence vs output price
Bubble = usage share. The line traces the Pareto frontier: models where nothing cheaper scores as high.
How to read: X-axis is list output price (log); Y-axis is benchmark score; bubble area is current token share. The connecting line joins models that dominate on price-for-score.
Benchmark score vs OpenRouter share — Intelligence
Top-right = high score AND heavily used · bottom-right = high score but under-adopted · top-left = popular despite lower score. Bubble size scales with usage share.
What a session actually costs in Hermes Agent
Median cost of one coding session by session length. 30-day window, log scale.
GPT-5.6 Sol Pro costs 4246× more per session than gpt-oss-20b in this agent.
How to read: Each row is a model. Dots run from short sessions to long ones — the further right, the more a median session costs. The number on the right is the 10–49 turns median.
- $5.87
- $4.30
- $2.18
- $1.68
- $1.67
- $1.21
- $1.19
- $1.01
- $0.981
- $0.829
- $0.797
- $0.736
Intelligence per dollar of real work
Benchmark score against the median 10–49 turns session cost in Hermes Agent.
Models on the dashed line are undominated: nothing is both cheaper and better-scoring. Bubble size is OpenRouter token share.
How to read: Cost runs left to right on a log scale, benchmark score bottom to top. Up and to the left is better value. Click a bubble for the model page.