Model Rankings
Rankings by capability, provenance and recency — what models accept, how much context they hold, who builds them and when they shipped. Capability facts are verified against each provider’s own documentation.
3,494 models tracked · 40 with verified capabilities · data as of 2026-08-13
Model ranking
30 models across every LiveBench category, with LMArena Elo alongside.
| Model | Overall | Reasoning | Coding | Agentic coding | Mathematics | Data analysis | Language | Instruction following | Elo |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 83.0 | 89.7 | 86.0 | 62.2 | 96.0 | 80.5 | 90.7 | 75.8 | 1506#1 |
| 2 | 81.1 | 91.7 | 83.9 | 56.2 | 96.2 | 79.8 | 87.7 | 71.8 | n/a |
| 3 | 80.2 | 89.7 | 82.1 | 54.0 | 95.9 | 81.6 | 87.4 | 70.7 | 1477#20 |
| 4 | 80.1 | 91.2 | 81.4 | 65.2 | 95.7 | 74.6 | 88.7 | 63.8 | n/a |
| 5 | 79.2 | 90.7 | 81.4 | 62.2 | 84.4 | 78.7 | 85.5 | 71.4 | n/a |
| 6 | 78.0 | 88.1 | 77.5 | 53.8 | 94.1 | 79.3 | 82.6 | 70.2 | 1465#39 |
| 7 | 78.0 | 90.0 | 77.5 | 57.6 | 91.2 | 76.5 | 78.6 | 74.3 | n/a |
| 8 | 77.9 | 90.6 | 78.2 | 54.9 | 94.9 | 79.3 | 82.9 | 64.6 | n/a |
| 9 | 77.6 | 91.2 | 75.0 | 54.2 | 93.1 | 75.3 | 82.7 | 71.5 | n/a |
| 10 | 77.0 | 84.0 | 76.5 | 44.1 | 91.0 | 78.5 | 85.4 | 79.1 | 1486#13 |
| 11 | 76.5 | 87.2 | 82.1 | 50.7 | 92.9 | 78.3 | 77.9 | 66.7 | 1494#7 |
| 12 | 76.2 | 89.2 | 81.8 | 50.5 | 94.3 | 66.0 | 79.7 | 72.0 | 1473#27 |
| 13 | 76.0 | 88.7 | 80.7 | 59.4 | 92.9 | 71.7 | 75.0 | 63.9 | n/a |
| 14 | 75.8 | 87.2 | 68.6 | 56.5 | 90.8 | 73.0 | 82.8 | 71.5 | 1469#35 |
| 15 | 75.3 | 87.7 | 77.2 | 58.5 | 87.1 | 72.5 | 74.3 | 69.6 | 1489#10 |
| 16 | 74.6 | 82.0 | 78.2 | 49.0 | 88.2 | 64.9 | 84.6 | 75.6 | n/a |
| 17 | 74.6 | 83.2 | 76.1 | 50.3 | 93.2 | 78.2 | 79.8 | 61.8 | 1435#83 |
| 18 | 74.5 | 88.7 | 78.2 | 49.0 | 89.3 | 69.9 | 83.3 | 63.3 | 1498#5 |
| 19 | 74.2 | 86.6 | 75.0 | 46.8 | 86.8 | 79.3 | 79.2 | 65.5 | n/a |
| 20 | 74.0 | 77.7 | 83.6 | 49.4 | 88.8 | 78.2 | 73.7 | 66.4 | n/a |
| 21 | 73.6 | 85.1 | 77.9 | 43.4 | 86.4 | 63.0 | 83.9 | 75.4 | 1484#15 |
| 22 | 73.6 | 85.6 | 82.9 | 48.4 | 87.2 | 78.0 | 72.6 | 60.1 | n/a |
| 23 | 73.2 | 78.6 | 79.7 | 51.8 | 89.8 | 73.7 | 76.2 | 62.3 | n/a |
| 24 | 73.1 | 83.3 | 74.2 | 43.6 | 85.2 | 71.8 | 79.7 | 74.0 | 1474#25 |
| 25 | 73.0 | 84.8 | 79.3 | 42.6 | 87.0 | 77.9 | 76.1 | 63.2 | 1472#30 |
| 26 | 72.6 | 80.1 | 79.7 | 39.7 | 90.4 | 74.4 | 81.3 | 62.5 | n/a |
| 27 | 71.9 | 78.3 | 71.0 | 49.4 | 88.4 | 72.8 | 73.5 | 70.1 | 1442#73 |
| 28 | 71.6 | 82.7 | 70.0 | 42.6 | 90.7 | 74.5 | 78.1 | 62.4 | 1458#50 |
| 29 | 70.5 | 79.4 | 78.6 | 46.9 | 84.3 | 65.1 | 75.1 | 64.4 | 1461#44 |
| 30 | 69.6 | 81.1 | 70.8 | 46.8 | 91.0 | 67.6 | 62.5 | 67.2 | n/a |
Two sources, never merged. Everything left of Elo is LiveBench, scored 0–100 and machine-marked; Elo is LMArena human preference. They share neither method nor scale, so agreement is real corroboration rather than one measurement counted twice. Shading marks the top five per column. The effort suffix matters — a max-effort run is a different measurement from the same model at low effort, so it is shown rather than hidden. LiveBench CC-BY-SA-4.0; LMArena CC-BY-4.0.
The frontier over time
Every arena-rated model by release date. Highlighted points held the record when they shipped.
21 of 120 rated models led the field on their release date. The vertical spread at any date is how much the field varies at one moment; the upward drift is progress.
Strength profile of the top models
Percentile within each arena category, for the ten highest-rated models.
| Model | Coding | Math | Creative writing | Instruction following | Hard prompts | Chinese | Japanese |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| Claude Opus 4.6 | 98 | 98 | 97 | 99 | 99 | 99 | 94 |
| Claude Opus 4.7 | 99 | 97 | 97 | 98 | 98 | 96 | 93 |
| Muse Spark 1.1 | 97 | 94 | 87 | 94 | 97 | 97 | 82 |
| Muse Spark | 96 | 84 | 95 | 89 | 96 | 91 | · |
| Gemini 3.1 Pro | 92 | 96 | 98 | 97 | 97 | 95 | 98 |
| Gemini 3 Pro | 90 | 91 | 99 | 92 | 94 | 92 | 99 |
| Gemini 3.6 Flash | 93 | 99 | 96 | 97 | 92 | 98 | 91 |
| GPT-5.5 | 87 | 97 | 90 | 93 | 92 | 96 | 96 |
| Qwen3.7-Max | 94 | 95 | 88 | 91 | 90 | 90 | · |
Percentile within each category. A dotted cell means not rated there — which is a different fact from ranking last, and is not shaded as if it were.
A single Elo number hides this. Two models a point apart overall can sit twenty percentiles apart on coding or on Japanese, which is the difference that matters when you are choosing one.
Popularity, open ecosystem
Hugging Face downloads over the last 30 days. Open weights only.
- Qwen3-0.6B27,629,624
- Qwen3-8B15,207,936
- Qwen3.5-9B12,423,514
- RoBERTa Base12,221,106
- Qwen2.5-1.5B12,108,126
- RoBERTa Large11,249,715
- Qwen2.5-7B11,218,582
- Gemma 4 26B A4B9,985,859
- Gemma 4 31B IT9,882,137
- Llama 3.2 1B9,198,251
A closed model has no repository, so GPT and Claude cannot appear here at all. This ranks the most-downloaded OPEN models, not the most-used models — there is no admissible free source for proprietary popularity, and inventing a proxy for it would be a confident number measuring nothing. Counts include automated pulls.
Capability coverage
Which input formats models actually accept, across the 40 verified against provider docs.
- text39
- image35
- video6
- file1
- audio1
This is the ranking that matters for content routing: if a format is rare, targeting it narrows your options sharply.
Context length
Distribution across models with a documented context window.
- 128K - 500K19
- 500K - 1M15
- Not stated6
Release velocity
Models released per quarter over the last three years.
Openness
Accessibility class, from Epoch AI. Models with no stated class are excluded rather than assumed closed.
- Open weights (unrestricted)30%
- Unreleased30%
- API access15%
- Open weights (restricted use)11%
- Open weights (non-commercial)8%
- Hosted access (no API)5%
Longest context
Verified context windows, largest first.
- 1GPT-5.4 · OpenAItext + image1,050,000
- 2GPT-5.4 Pro · OpenAItext + image1,050,000
- 3GPT-5.5 · OpenAItext + image1,050,000
- 4GPT-5.5 Pro · OpenAItext + image1,050,000
- 5GPT-5.6 Luna · OpenAItext + image1,050,000
- 6GPT-5.6 Sol · OpenAItext + image1,050,000
- 7GPT-5.6 Terra · OpenAItext + image1,050,000
- 8GLM-5.2 · Z.ai (Zhipu AI)text1,048,576
- 9Nemotron 3.5 Lightning · NVIDIAtext1,048,576
- 10Claude Fable 5 · Anthropictext + image1,000,000
Newest models
Most recently released, from the Epoch AI registry.
- 1Grok 4.6 · xAI2026-08-12
- 2GPT-5.5 Cyber · OpenAI2026-08-11
- 3GPT-5.6 Cyber · OpenAI2026-08-11
- 4Nemotron 3.5 Lightning · NVIDIA2026-08-11
- 5Muse Glimmer · Meta AI2026-08-10
- 6Motif-3 · Motif Technologies2026-08-07
- 7Muse Spark 1.2 · Meta AI2026-08-05
- 8DeepSeek V4 Flash 0731 · DeepSeek2026-07-31
- 9K-EXAONE 2.0 · LG AI Research2026-07-31
- 10Gemini Robotics ER 2 · Google DeepMind2026-07-30
Speed and availability
Tokens per second, time to first token, and derived uptime.
Needs: LBOX first-party probes, blocked on the provider key inventory (gap OG-1.4). Provider status feeds were disabled under v11 §24 pending a per-vendor terms read. Artificial Analysis measures exactly this and is NOT the answer: its free tier is internal-use-only, so its numbers can inform the routing engine and can never appear here (§26.4).
Browse and filter the full registry on Models.