Model Rankings

Rankings by capability, provenance and recency — what models accept, how much context they hold, who builds them and when they shipped. Capability facts are verified against each provider’s own documentation.

3,494 models tracked · 40 with verified capabilities · data as of 2026-08-13

Model ranking

30 models across every LiveBench category, with LMArena Elo alongside.

ModelOverallReasoningCodingAgentic codingMathematicsData analysisLanguageInstruction followingElo
1ACClaude Fable 583.089.786.062.296.080.590.775.81506#1
2OIGPT-5.6 Sol81.191.783.956.296.279.887.771.8n/a
3OIGPT-5.580.289.782.154.095.981.687.470.71477#20
4ACClaude Opus 580.191.281.465.295.774.688.763.8n/a
5MTKimi K379.290.781.462.284.478.785.571.4n/a
6OIGPT-5.478.088.177.553.894.179.382.670.21465#39
7MAMuse Spark 1.278.090.077.557.691.276.578.674.3n/a
8OIGPT-5.6 Terra77.990.678.254.994.979.382.964.6n/a
9XIGrok 4.677.691.275.054.293.175.382.771.5n/a
10GDGemini 3.1 Pro77.084.076.544.191.078.585.479.11486#13
11ACClaude Opus 4.776.587.282.150.792.978.377.966.71494#7
12ACClaude Opus 4.876.289.281.850.594.366.079.772.01473#27
13ACClaude Sonnet 576.088.780.759.492.971.775.063.9n/a
14XIGrok 4.575.887.268.656.590.873.082.871.51469#35
15MAMuse Spark 1.175.387.777.258.587.172.574.369.61489#10
16GEGemini 3.5 Flash74.682.078.249.088.264.984.675.6n/a
17OIGPT-5.274.683.276.150.393.278.279.861.81435#83
18ACClaude Opus 4.674.588.778.249.089.369.983.363.31498#5
19DKDeepSeek V4 Flash 073174.286.675.046.886.879.379.265.5n/a
20OIGPT-5.2 Codex74.077.783.649.488.878.273.766.4n/a
21GDGemini 3.6 Flash73.685.177.943.486.463.083.975.41484#15
22OIGPT-5.6 Luna73.685.682.948.487.278.072.660.1n/a
23Z(GLM-5.273.278.679.751.889.873.776.262.3n/a
24AAQwen3.7-Max73.183.374.243.685.271.879.774.01474#25
25ACClaude Sonnet 4.673.084.879.342.687.077.976.163.21472#30
26ACClaude Opus 4.572.680.179.739.790.474.481.362.5n/a
27TMInkling71.978.371.049.488.472.873.570.11442#73
28DKDeepSeek-V4-Pro71.682.770.042.690.774.578.162.41458#50
29MTKimi K2.670.579.478.646.984.365.175.164.41461#44
30OIGPT-5.4 Nano69.681.170.846.891.067.662.567.2n/a

Two sources, never merged. Everything left of Elo is LiveBench, scored 0–100 and machine-marked; Elo is LMArena human preference. They share neither method nor scale, so agreement is real corroboration rather than one measurement counted twice. Shading marks the top five per column. The effort suffix matters — a max-effort run is a different measurement from the same model at low effort, so it is shown rather than hidden. LiveBench CC-BY-SA-4.0; LMArena CC-BY-4.0.

The frontier over time

Every arena-rated model by release date. Highlighted points held the record when they shipped.

93110861240139515492023202420252026LLaMA-13B · 2023-02-27 — 974WizardLM 70B · 2023-04-24 — 1184PaLM 2 · 2023-05-10 — 1138Llama 2-7B · 2023-07-18 — 1107Llama 2-13B · 2023-07-18 — 1141Llama 2-70B · 2023-07-18 — 1170Qwen-14B · 2023-09-28 — 1139Mistral 7B · 2023-10-10 — 1110Mistral Medium · 2023-12-11 — 1222DeepSeek LLM 67B · 2024-01-05 — 1184OLMo-7B · 2024-02-01 — 1073Qwen1.5-7B · 2024-02-04 — 1143Qwen1.5-14B · 2024-02-04 — 1191Qwen1.5-72B · 2024-02-04 — 1233Qwen1.5-32B · 2024-02-05 — 1203Gemma 2B · 2024-02-21 — 1093Gemma 7B · 2024-02-21 — 1137Gemma 1.1 7B Instruct · 2024-02-24 — 1182Command R+ · 2024-04-04 — 1226Llama 3-8B · 2024-04-18 — 1223Llama 3-70B · 2024-04-18 — 1276Qwen1.5-110B · 2024-04-25 — 1234GLM-4 (0520) · 2024-05-20 — 1273Qwen2-72B · 2024-06-07 — 1261Nemotron-4 340B · 2024-06-14 — 1277Gemma 2 2B · 2024-06-24 — 1200Gemma 2 9B · 2024-06-24 — 1267Gemma 2 27B · 2024-06-24 — 1289Llama 3.1-8B · 2024-07-23 — 1211Llama 3.1-70B · 2024-07-23 — 1293Llama 3.1-405B · 2024-07-23 — 1333GLM-4-Plus · 2024-08-29 — 1319DeepSeek-V2.5 · 2024-09-06 — 1307o1-mini · 2024-09-12 — 1337Qwen2.5-72B · 2024-09-19 — 1303Llama 3.2 1B · 2024-09-24 — 1111Llama 3.2 3B · 2024-09-24 — 1166Llama-3.1-Nemotron-70B-Instruct · 2024-10-12 — 1299Granite 3.0 2B · 2024-10-21 — 1156Granite 3.0 8B · 2024-10-21 — 1182Qwen2.5-Coder (32B) · 2024-11-12 — 1270Llama 3.3 70B · 2024-12-06 — 1318Phi-4 · 2024-12-12 — 1256Granite 3.1 2B · 2024-12-18 — 1178Granite 3.1 8B · 2024-12-18 — 1208DeepSeek-V3 · 2024-12-24 — 1359DeepSeek-R1 · 2025-01-20 — 1398Qwen2.5-Max · 2025-01-28 — 1374o3-mini · 2025-01-31 — 1348Grok-3 mini · 2025-02-19 — 1356Mercury · 2025-02-27 — 1306QwQ-32B · 2025-03-06 — 1154Gemma 3 4B · 2025-03-12 — 1303Gemma 3 12B · 2025-03-12 — 1342Gemma 3 27B · 2025-03-12 — 1366Gemini 2.5 Flash · 2025-04-17 — 1410Qwen3-30B-A3B · 2025-04-29 — 1327Qwen3-32B · 2025-04-29 — 1347Qwen3-235B-A22B · 2025-04-29 — 1375Qwen3-Coder-480B-A35B · 2025-07-22 — 1388gpt-oss-20b · 2025-08-05 — 1318gpt-oss-120b · 2025-08-05 — 1352GLM-4.5-Air · 2025-08-05 — 1373GLM-4.5 · 2025-08-05 — 1411GPT-5 · 2025-08-07 — 1427GLM-4.5V · 2025-08-15 — 1353DeepSeek-V3.1 · 2025-08-21 — 1417LongCat-Flash · 2025-09-01 — 1401Qwen3-Max · 2025-09-05 — 1434Qwen3-Next-80B-A3B · 2025-09-10 — 1401Grok 4 Fast · 2025-09-19 — 1420DeepSeek-V3.1-Terminus · 2025-09-22 — 1415GLM-4.6 · 2025-09-30 — 1424MiniMax-M2 · 2025-10-27 — 1345GPT-5.1 · 2025-11-13 — 1439Grok 4.1 · 2025-11-17 — 1459Gemini 3 Pro · 2025-11-18 — 1485Olmo 3.1 32B Think · 2025-11-20 — 1285Mistral Large 3 · 2025-12-02 — 1415GPT-5.2 · 2025-12-11 — 1435Gemini 3 Flash · 2025-12-17 — 1473GLM 4.7 Flash · 2025-12-22 — 1367GLM-4.7 · 2025-12-22 — 1442MiniMax-M2.1 · 2025-12-23 — 1384Step 3.5 Flash · 2026-02-02 — 1395Claude Opus 4.6 · 2026-02-05 — 1498GLM-5 · 2026-02-11 — 1457MiniMax-M2.5 · 2026-02-12 — 1390Qwen3.5 397B-A17B · 2026-02-13 — 1442Claude Sonnet 4.6 · 2026-02-17 — 1472Gemini 3.1 Pro · 2026-02-19 — 1486Qwen3.5-35B-A3B · 2026-02-24 — 1396Qwen3.5-27B · 2026-02-24 — 1408Qwen3.5-122B-A10B · 2026-02-24 — 1417Gemini 3.1 Flash-Lite · 2026-03-03 — 1432GPT-5.4 · 2026-03-05 — 1465MiniMax-M2.7 · 2026-03-18 — 1416MiMo-V2-Pro · 2026-03-18 — 1448Gemma 4 26B A4B · 2026-04-02 — 1438Gemma 4 31B IT · 2026-04-02 — 1451GLM-5.1 · 2026-04-07 — 1467Muse Spark · 2026-04-08 — 1488Claude Opus 4.7 · 2026-04-16 — 1494Grok 4.3 Beta · 2026-04-17 — 1442Kimi K2.6 · 2026-04-20 — 1461MiMo-V2.5-Pro · 2026-04-23 — 1468GPT-5.5 · 2026-04-23 — 1477DeepSeek-V4-Flash · 2026-04-24 — 1435DeepSeek-V4-Pro · 2026-04-24 — 1458Mistral Medium 3.5 · 2026-04-29 — 1427Qwen3.7-Max · 2026-05-19 — 1474Claude Opus 4.8 · 2026-05-28 — 1473MiniMax-M3 · 2026-06-01 — 1443Claude Fable 5 · 2026-06-09 — 1506Grok 4.5 · 2026-07-08 — 1469Muse Spark 1.1 · 2026-07-09 — 1489Inkling · 2026-07-15 — 1442Gemini 3.5 Flash-Lite · 2026-07-21 — 1459Gemini 3.6 Flash · 2026-07-21 — 1484Muse Glimmer · 2026-08-10 — 1426release date

21 of 120 rated models led the field on their release date. The vertical spread at any date is how much the field varies at one moment; the upward drift is progress.

Strength profile of the top models

Percentile within each arena category, for the ten highest-rated models.

ModelCodingMathCreative writingInstruction followingHard promptsChineseJapanese
Claude Fable 5100100100100100100100
Claude Opus 4.698989799999994
Claude Opus 4.799979798989693
Muse Spark 1.197948794979782
Muse Spark968495899691·
Gemini 3.1 Pro92969897979598
Gemini 3 Pro90919992949299
Gemini 3.6 Flash93999697929891
GPT-5.587979093929696
Qwen3.7-Max949588919090·

Percentile within each category. A dotted cell means not rated there — which is a different fact from ranking last, and is not shaded as if it were.

A single Elo number hides this. Two models a point apart overall can sit twenty percentiles apart on coding or on Japanese, which is the difference that matters when you are choosing one.

Popularity, open ecosystem

Hugging Face downloads over the last 30 days. Open weights only.

  • Qwen3-0.6B27,629,624
  • Qwen3-8B15,207,936
  • Qwen3.5-9B12,423,514
  • RoBERTa Base12,221,106
  • Qwen2.5-1.5B12,108,126
  • RoBERTa Large11,249,715
  • Qwen2.5-7B11,218,582
  • Gemma 4 26B A4B9,985,859
  • Gemma 4 31B IT9,882,137
  • Llama 3.2 1B9,198,251

A closed model has no repository, so GPT and Claude cannot appear here at all. This ranks the most-downloaded OPEN models, not the most-used models — there is no admissible free source for proprietary popularity, and inventing a proxy for it would be a confident number measuring nothing. Counts include automated pulls.

Capability coverage

Which input formats models actually accept, across the 40 verified against provider docs.

  • text39
  • image35
  • video6
  • file1
  • audio1

This is the ranking that matters for content routing: if a format is rare, targeting it narrows your options sharply.

Context length

Distribution across models with a documented context window.

  • 128K - 500K19
  • 500K - 1M15
  • Not stated6

Release velocity

Models released per quarter over the last three years.

2962023 Q3: 712023 Q4: 1762024 Q1: 1572024 Q2: 2392024 Q3: 2962024 Q4: 2222025 Q1: 1832025 Q2: 1822025 Q3: 1162025 Q4: 862026 Q1: 402026 Q2: 462026 Q3: 292023 Q32024 Q22025 Q12025 Q42026 Q3

Openness

Accessibility class, from Epoch AI. Models with no stated class are excluded rather than assumed closed.

Open weights (unrestricted): 825 (30%)Unreleased: 823 (30%)API access: 414 (15%)Open weights (restricted use): 294 (11%)Open weights (non-commercial): 228 (8%)Hosted access (no API): 122 (5%)2,706models
  • Open weights (unrestricted)30%
  • Unreleased30%
  • API access15%
  • Open weights (restricted use)11%
  • Open weights (non-commercial)8%
  • Hosted access (no API)5%

Longest context

Verified context windows, largest first.

  1. 1GPT-5.4 · OpenAItext + image1,050,000
  2. 2GPT-5.4 Pro · OpenAItext + image1,050,000
  3. 3GPT-5.5 · OpenAItext + image1,050,000
  4. 4GPT-5.5 Pro · OpenAItext + image1,050,000
  5. 5GPT-5.6 Luna · OpenAItext + image1,050,000
  6. 6GPT-5.6 Sol · OpenAItext + image1,050,000
  7. 7GPT-5.6 Terra · OpenAItext + image1,050,000
  8. 8GLM-5.2 · Z.ai (Zhipu AI)text1,048,576
  9. 9Nemotron 3.5 Lightning · NVIDIAtext1,048,576
  10. 10Claude Fable 5 · Anthropictext + image1,000,000

Newest models

Most recently released, from the Epoch AI registry.

  1. 1Grok 4.6 · xAI2026-08-12
  2. 2GPT-5.5 Cyber · OpenAI2026-08-11
  3. 3GPT-5.6 Cyber · OpenAI2026-08-11
  4. 4Nemotron 3.5 Lightning · NVIDIA2026-08-11
  5. 5Muse Glimmer · Meta AI2026-08-10
  6. 6Motif-3 · Motif Technologies2026-08-07
  7. 7Muse Spark 1.2 · Meta AI2026-08-05
  8. 8DeepSeek V4 Flash 0731 · DeepSeek2026-07-31
  9. 9K-EXAONE 2.0 · LG AI Research2026-07-31
  10. 10Gemini Robotics ER 2 · Google DeepMind2026-07-30

Speed and availability

Tokens per second, time to first token, and derived uptime.

Needs: LBOX first-party probes, blocked on the provider key inventory (gap OG-1.4). Provider status feeds were disabled under v11 §24 pending a per-vendor terms read. Artificial Analysis measures exactly this and is NOT the answer: its free tier is internal-use-only, so its numbers can inform the routing engine and can never appear here (§26.4).

Browse and filter the full registry on Models.