SOTA Model — the live state-of-the-art AI leaderboard. Current SOTA: Anthropic: Claude Fable 5.1 by Anthropic.
Current state of the art · v4.0 · streaming
“Claude Fable 5.1 excels in SWE-bench Verified, Aider Polyglot, Terminal-Bench, Chatbot Arena, LiveBench, and long-horizon software-engineering evaluations.”
SOTA Elo
+9 vs #2 in the field
SOTA GPQA
Graduate-level science QA
Max Context
Token window at the top
Avg Output $
Across tracked frontier tier
Pick a lens
SOTA depends on the job.
The general SOTA above answers "what's the best model." These leaderboards answer "best for what."
Model Telemetry
| 01 | Anthropic: Claude Fable 5.1Anthropic SOTA | 1638 | 70 | 85.9 | 90.7 | 86.1 | 67.5 | 28.9 | 85.4 | 1M | $50.00 | 16 days ago |
| 02 | Sakana: Fugu Ultra v2Sakana | 1629 | 69 | 85.0 | 90.1 | 83.6 | 70.4 | 33.1 | 81.8 | 1M | $30.00 | 6 days ago |
| 03 | Qwen: Qwen3.8 Max (0902)Alibaba | 1625 | 68 | 84.3 | 89.4 | 81.9 | 65.8 | 27.4 | 81.2 | 1M | $6.00 | 14 days ago |
| 04 | 1621 | 68 | 84.0 | 89.8 | 81.2 | 66.3 | 26.8 | 80.7 | 500K | $6.00 | 1 month ago | |
| 05 | MoonshotAI: Kimi K3Moonshot | 1615 | 67 | 82.9 | 88.6 | 80.8 | 64.6 | 25.7 | 82.0 | 1.0M | $15.00 | 2 months ago |
| 06 | Z.ai: GLM 5.3Z.ai | 1609 | 66 | 82.5 | 87.9 | 80.1 | 63.9 | 24.5 | 81.5 | 1.3M | $4.40 | 1 month ago |
| 07 | 1604 | 65 | 81.6 | 87.2 | 78.8 | 62.1 | 25.9 | 78.5 | 1.0M | $4.25 | 15 days ago | |
| 08 | OpenAI: GPT-6 Astra ProOpenAI | 1600 | 70 | 86.0 | 91.1 | 85.3 | 70.1 | 32.0 | 83.8 | 1.1M | $50.00 | 13 days ago |
| 09 | DeepSeek: DeepSeek V4 Pro 0813DeepSeek | 1596 | 64 | 80.8 | 86.5 | 78.6 | 61.7 | 23.1 | 79.4 | 1.0M | $1.74 | 1 month ago |
| 10 | Anthropic: Claude Opus 5Anthropic | 1591 | 64 | 81.1 | 86.8 | 79.6 | 61.2 | 23.8 | 81.1 | 1M | $25.00 | 1 month ago |
| 11 | OpenAI: GPT-5.6 SolOpenAI | 1587 | 64 | 80.9 | 86.2 | 78.3 | 63.2 | 24.1 | 78.7 | 1.1M | $10.00 | 2 months ago |
| 12 | Qwen: Qwen3.8 2.4T A95BAlibaba | 1578 | 62 | 79.8 | 85.6 | 76.9 | 60.1 | 22.8 | 78.3 | 1.0M | $6.00 | 1 month ago |
| 13 | Thinking Machines: InklingThinking Machines | 1574 | 61 | 79.1 | 84.7 | 76.4 | 59.4 | 22.4 | 77.6 | 1.0M | $4.05 | 2 months ago |
| 14 | Anthropic: Claude Fable 5Anthropic | 1569 | 61 | 79.3 | 84.9 | 77.2 | 59.7 | 21.9 | 79.0 | 1M | $50.00 | 3 months ago |
| 15 | Sakana: Fugu UltraSakana | 1565 | 60 | 78.7 | 84.3 | 75.8 | 61.0 | 25.1 | 75.6 | 1M | $30.00 | 2 months ago |
| 16 | 1561 | 60 | 78.6 | 84.6 | 75.5 | 59.8 | 21.7 | 75.9 | 500K | $6.00 | 2 months ago | |
| 17 | OpenAI: GPT-5.5 ProOpenAI | 1557 | 60 | 79.0 | 85.1 | 76.7 | 61.8 | 23.5 | 76.8 | 1.1M | $180.00 | 4 months ago |
| 18 | Anthropic: Claude Opus 4.8Anthropic | 1553 | 59 | 77.9 | 83.8 | 76.1 | 58.5 | 20.4 | 78.1 | 1M | $25.00 | 3 months ago |
| 19 | Z.ai: GLM 5.2Z.ai | 1548 | 58 | 77.4 | 83.4 | 74.9 | 58.1 | 20.0 | 76.4 | 1.0M | $4.40 | 3 months ago |
| 20 | MoonshotAI: Kimi K2.7 CodeMoonshot | 1544 | 58 | 76.8 | 82.5 | 75.6 | 57.7 | 18.9 | 79.2 | 262K | $3.21 | 3 months ago |
| 21 | 1540 | 57 | 76.6 | 82.8 | 74.5 | 56.9 | 20.7 | 74.3 | 1.0M | $4.25 | 1 month ago | |
| 22 | NVIDIA: Nemotron 3 UltraNVIDIA | 1535 | 56 | 75.7 | 81.9 | 73.1 | 56.4 | 18.5 | 74.6 | 262K | $3.13 | 3 months ago |
| 23 | Qwen: Qwen3.7 MaxAlibaba | 1531 | 55 | 75.3 | 81.5 | 72.8 | 55.7 | 18.1 | 73.9 | 1M | $4.42 | 3 months ago |
| 24 | Anthropic: Claude Haiku 4.5Anthropic | 1548 | 60 | 83.2 | 83.5 | 68.9 | 50.8 | 17.9 | 78.6 | 200K | $5.00 | 11 months ago |
| 25 | Qwen: Qwen3.6 PlusAlibaba | — | — | — | — | — | — | — | — | 1M | $1.95 | 5 months ago |
| 26 | 1489 | 56 | 73.7 | 71.6 | 64.8 | 51.6 | 11.9 | 74.1 | 131K | $1.10 | 1 month ago | |
| 27 | Google WeatherNext 3Google | — | — | — | — | — | — | — | — | — | — | — |
| 28 | Shieldstral 1.0 3BMistral AI | — | — | — | — | — | — | — | — | — | — | — |
| 29 | MiniMax H3MiniMax | — | — | — | — | — | — | — | — | — | — | — |
| 30 | Qwen 3.8Alibaba / Qwen | — | — | — | — | — | — | — | — | — | — | — |
| 31 | GPT-transcribeOpenAI | — | — | — | — | — | — | — | — | — | — | — |
| 32 | GPT-5.6 TerraOpenAI | 1462 | — | — | 80.2 | — | — | — | — | 400K | $4.80 | 2 months ago |
| 33 | Gemini 3.5 Flash CyberGoogle | — | — | — | — | — | — | — | — | — | — | — |
| 34 | Gemini 3.6 FlashGoogle | — | — | — | — | — | — | — | — | — | — | — |
| 35 | GPT-Live-1OpenAI | — | — | — | — | — | — | — | — | — | — | — |
| 36 | Claude Mythos 5Anthropic | — | — | — | — | — | — | — | — | — | — | — |
| 37 | Thinking Machines Inkling-Small 276B-A12BThinking Machines | — | — | — | — | — | — | — | — | — | — | — |
| 38 | Gemini Robotics 2.0Google DeepMind | — | — | — | — | — | — | — | — | — | — | — |
| 39 | Qwen3.8-MaxAlibaba | — | — | — | — | — | — | — | — | — | — | — |
| 40 | GPT-5.6 LunaOpenAI | 1421 | — | — | 74.9 | — | — | — | — | 200K | $0.80 | 2 months ago |
| 41 | Claude 4.5 OpusAnthropic | 1458 | — | — | 81.9 | — | — | — | — | 500K | $24.00 | 4 months ago |
| 42 | OpenAI: GPT-5.6 LunaOpenAI | — | — | — | — | — | — | — | — | 1.1M | $0.60 | 2 months ago |
| 43 | — | — | — | — | — | — | — | — | 131K | $6.00 | 2 months ago | |
| 44 | Qwen: Qwen3.8 MaxAlibaba | 1609 | 68 | 89.4 | 91.7 | 78.5 | 70.3 | 54.7 | 84.3 | 1M | $6.00 | 1 month ago |
| 45 | — | — | — | — | — | — | — | — | — | — | — | |
| 46 | Nex AGI Nex-N2.5 MaxNex AGI | — | — | — | — | — | — | — | — | — | — | — |
| 47 | Qwen 3 MaxAlibaba | 1371 | — | — | 68.9 | — | — | — | — | 262K | $1.50 | 3 months ago |
| 48 | Mistral Large 3Mistral | 1358 | — | — | 66.4 | — | — | — | — | 256K | $6.00 | 4 months ago |
| 49 | Qwen3.8-MaxAlibaba | — | — | — | — | — | — | — | — | — | — | — |
| 50 | — | — | — | — | — | — | — | — | — | — | — | |
| 51 | DeepSeek: DeepSeek V4 Flash 0731DeepSeek | 1495 | 60 | 75.0 | 80.0 | 70.0 | 57.0 | 17.0 | 81.0 | 1.0M | $0.18 | 1 month ago |
| 52 | — | — | — | — | — | — | — | — | — | — | — | |
| 53 | Grok 4xAI | 1402 | — | — | 74.1 | — | — | — | — | 256K | $8.00 | 3 months ago |
| 54 | DeepSeek V4DeepSeek | 1381 | — | — | 69.8 | — | — | — | — | 128K | $0.28 | 3 months ago |
| 55 | GPT-LiveOpenAI | — | — | — | — | — | — | — | — | — | — | — |
| 56 | Gemini Robotics 2Google DeepMind | — | — | — | — | — | — | — | — | — | — | — |
| 57 | Qwen3.8-MaxAlibaba / Qwen | — | — | — | — | — | — | — | — | — | — | — |
| 58 | Grok 4.6xAI | — | — | — | — | — | — | — | — | — | — | — |
| 59 | Claude 4.5 SonnetAnthropic | 1418 | — | — | 76.8 | — | — | — | — | 400K | $7.20 | 4 months ago |
| 60 | Qwen: Qwen3.7 PlusAlibaba | 1484 | 64 | 73.5 | 80.1 | 64.7 | 51.9 | 24.0 | 73.0 | 1M | $1.28 | 3 months ago |
| 61 | Llama 4 405BMeta | 1389 | — | — | 71.2 | — | — | — | — | 128K | $2.70 | 4 months ago |
| 62 | Opus 5Anthropic | — | — | — | — | — | — | — | — | — | — | — |
| 63 | OpenAI: GPT-5.5OpenAI | 1546 | 59 | 82.9 | 83.8 | 70.4 | 65.8 | 36.2 | 73.9 | 1.1M | $30.00 | 4 months ago |
| 64 | 1499 | 43 | 74.6 | 83.1 | 68.2 | 57.2 | 11.4 | 75.1 | 2M | $2.50 | 5 months ago | |
| 65 | Google: Gemini 3.8 FlashGoogle | 1558 | 63 | 80.8 | 82.5 | 71.5 | 59.8 | 24.6 | 74.7 | 1.0M | $3.75 | 15 days ago |
| 66 | Claude MythosAnthropic | — | — | — | — | — | — | — | — | — | — | — |
| 67 | Google TimesFM 3Google | — | — | — | — | — | — | — | — | — | — | — |
| 68 | SIMA 2Google DeepMind | — | — | — | — | — | — | — | — | — | — | — |
| 69 | Google: Gemini 3.6 FlashGoogle | 1498 | 55 | 75.1 | 80.9 | 70.8 | 57.2 | 14.1 | 75.5 | 1.0M | $3.75 | 1 month ago |
| 70 | Claude Opus 5 (Fast)Anthropic | 1629 | 69 | 84.8 | 90.5 | 80.5 | 66.8 | 35.4 | 88.0 | 1M | $50.00 | 1 month ago |
| 71 | — | — | — | — | — | — | — | — | 1.0M | $1.50 | 4 months ago | |
| 72 | — | — | — | — | — | — | — | — | — | — | — | |
| 73 | Mistral: Devstral 2 2512Mistral | 1518 | 54 | 72.4 | 80.7 | 71.4 | 58.1 | 18.9 | 82.1 | 262K | $2.00 | 9 months ago |
| 74 | — | — | — | — | — | — | — | — | 131K | $3.00 | 3 months ago | |
| 75 | DeepSeek: DeepSeek V4 Pro 0423DeepSeek | 1568 | 64 | 76.8 | 83.5 | 69.7 | 58.2 | 21.8 | 80.1 | 1.0M | $3.20 | 4 months ago |
| 76 | Qwen 3.8Alibaba | — | — | — | — | — | — | — | — | — | — | — |
| 77 | Anthropic: Claude Opus 4.8 (Fast)Anthropic | 1518 | 49 | 78.8 | 86.7 | 72.8 | 60.9 | 14.3 | 79.8 | 1M | $50.00 | 3 months ago |
| 78 | Claude Mythos 5.1Anthropic | — | — | — | — | — | — | — | — | — | — | — |
| 79 | Qwen Drive 1.0Alibaba Qwen | — | — | — | — | — | — | — | — | — | — | — |
| 80 | 1478 | 62 | 73.5 | 81.9 | 63.5 | 52.4 | 19.8 | 77.4 | 256K | $2.00 | 4 months ago |
Anthropic · 16 days ago
Elo
1638
AA
70
GPQA
90.7
SWE
86
Term
68
ARC
29
LiveB
86
Ctx
1M
Sakana · 6 days ago
Elo
1629
AA
69
GPQA
90.1
SWE
84
Term
70
ARC
33
LiveB
85
Ctx
1M
Alibaba · 14 days ago
Elo
1625
AA
68
GPQA
89.4
SWE
82
Term
66
ARC
27
LiveB
84
Ctx
1M
xAI · 1 month ago
Elo
1621
AA
68
GPQA
89.8
SWE
81
Term
66
ARC
27
LiveB
84
Ctx
500K
Moonshot · 2 months ago
Elo
1615
AA
67
GPQA
88.6
SWE
81
Term
65
ARC
26
LiveB
83
Ctx
1.0M
Z.ai · 1 month ago
Elo
1609
AA
66
GPQA
87.9
SWE
80
Term
64
ARC
25
LiveB
83
Ctx
1.3M
Meta · 15 days ago
Elo
1604
AA
65
GPQA
87.2
SWE
79
Term
62
ARC
26
LiveB
82
Ctx
1.0M
OpenAI · 13 days ago
Elo
1600
AA
70
GPQA
91.1
SWE
85
Term
70
ARC
32
LiveB
86
Ctx
1.1M
DeepSeek · 1 month ago
Elo
1596
AA
64
GPQA
86.5
SWE
79
Term
62
ARC
23
LiveB
81
Ctx
1.0M
Anthropic · 1 month ago
Elo
1591
AA
64
GPQA
86.8
SWE
80
Term
61
ARC
24
LiveB
81
Ctx
1M
OpenAI · 2 months ago
Elo
1587
AA
64
GPQA
86.2
SWE
78
Term
63
ARC
24
LiveB
81
Ctx
1.1M
Alibaba · 1 month ago
Elo
1578
AA
62
GPQA
85.6
SWE
77
Term
60
ARC
23
LiveB
80
Ctx
1.0M
Thinking Machines · 2 months ago
Elo
1574
AA
61
GPQA
84.7
SWE
76
Term
59
ARC
22
LiveB
79
Ctx
1.0M
Anthropic · 3 months ago
Elo
1569
AA
61
GPQA
84.9
SWE
77
Term
60
ARC
22
LiveB
79
Ctx
1M
Sakana · 2 months ago
Elo
1565
AA
60
GPQA
84.3
SWE
76
Term
61
ARC
25
LiveB
79
Ctx
1M
xAI · 2 months ago
Elo
1561
AA
60
GPQA
84.6
SWE
76
Term
60
ARC
22
LiveB
79
Ctx
500K
OpenAI · 4 months ago
Elo
1557
AA
60
GPQA
85.1
SWE
77
Term
62
ARC
24
LiveB
79
Ctx
1.1M
Anthropic · 3 months ago
Elo
1553
AA
59
GPQA
83.8
SWE
76
Term
59
ARC
20
LiveB
78
Ctx
1M
Z.ai · 3 months ago
Elo
1548
AA
58
GPQA
83.4
SWE
75
Term
58
ARC
20
LiveB
77
Ctx
1.0M
Moonshot · 3 months ago
Elo
1544
AA
58
GPQA
82.5
SWE
76
Term
58
ARC
19
LiveB
77
Ctx
262K
Meta · 1 month ago
Elo
1540
AA
57
GPQA
82.8
SWE
75
Term
57
ARC
21
LiveB
77
Ctx
1.0M
NVIDIA · 3 months ago
Elo
1535
AA
56
GPQA
81.9
SWE
73
Term
56
ARC
19
LiveB
76
Ctx
262K
Alibaba · 3 months ago
Elo
1531
AA
55
GPQA
81.5
SWE
73
Term
56
ARC
18
LiveB
75
Ctx
1M
Anthropic · 11 months ago
Elo
1548
AA
60
GPQA
83.5
SWE
69
Term
51
ARC
18
LiveB
83
Ctx
200K
Alibaba · 5 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
1M
Meta · 1 month ago
Elo
1489
AA
56
GPQA
71.6
SWE
65
Term
52
ARC
12
LiveB
74
Ctx
131K
Google · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Mistral AI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
MiniMax · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Alibaba / Qwen · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
OpenAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
OpenAI · 2 months ago
Elo
1462
AA
—
GPQA
80.2
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
400K
Google · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Google · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
OpenAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Anthropic · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Thinking Machines · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Google DeepMind · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Alibaba · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
OpenAI · 2 months ago
Elo
1421
AA
—
GPQA
74.9
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
200K
Anthropic · 4 months ago
Elo
1458
AA
—
GPQA
81.9
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
500K
OpenAI · 2 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
1.1M
Aion · 2 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
131K
Alibaba · 1 month ago
Elo
1609
AA
68
GPQA
91.7
SWE
79
Term
70
ARC
55
LiveB
89
Ctx
1M
xAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Nex AGI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Alibaba · 3 months ago
Elo
1371
AA
—
GPQA
68.9
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
262K
Mistral · 4 months ago
Elo
1358
AA
—
GPQA
66.4
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
256K
Alibaba · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
NVIDIA · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
DeepSeek · 1 month ago
Elo
1495
AA
60
GPQA
80.0
SWE
70
Term
57
ARC
17
LiveB
75
Ctx
1.0M
Z.ai · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
xAI · 3 months ago
Elo
1402
AA
—
GPQA
74.1
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
256K
DeepSeek · 3 months ago
Elo
1381
AA
—
GPQA
69.8
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
128K
OpenAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Google DeepMind · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Alibaba / Qwen · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
xAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Anthropic · 4 months ago
Elo
1418
AA
—
GPQA
76.8
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
400K
Alibaba · 3 months ago
Elo
1484
AA
64
GPQA
80.1
SWE
65
Term
52
ARC
24
LiveB
74
Ctx
1M
Meta · 4 months ago
Elo
1389
AA
—
GPQA
71.2
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
128K
Anthropic · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
OpenAI · 4 months ago
Elo
1546
AA
59
GPQA
83.8
SWE
70
Term
66
ARC
36
LiveB
83
Ctx
1.1M
xAI · 5 months ago
Elo
1499
AA
43
GPQA
83.1
SWE
68
Term
57
ARC
11
LiveB
75
Ctx
2M
Google · 15 days ago
Elo
1558
AA
63
GPQA
82.5
SWE
72
Term
60
ARC
25
LiveB
81
Ctx
1.0M
Anthropic · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Google · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Google DeepMind · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Google · 1 month ago
Elo
1498
AA
55
GPQA
80.9
SWE
71
Term
57
ARC
14
LiveB
75
Ctx
1.0M
Anthropic · 1 month ago
Elo
1629
AA
69
GPQA
90.5
SWE
81
Term
67
ARC
35
LiveB
85
Ctx
1M
Google · 4 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
1.0M
xAI · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Mistral · 9 months ago
Elo
1518
AA
54
GPQA
80.7
SWE
71
Term
58
ARC
19
LiveB
72
Ctx
262K
Google · 3 months ago
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
131K
DeepSeek · 4 months ago
Elo
1568
AA
64
GPQA
83.5
SWE
70
Term
58
ARC
22
LiveB
77
Ctx
1.0M
Alibaba · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Anthropic · 3 months ago
Elo
1518
AA
49
GPQA
86.7
SWE
73
Term
61
ARC
14
LiveB
79
Ctx
1M
Anthropic · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
Alibaba Qwen · —
Elo
—
AA
—
GPQA
—
SWE
—
Term
—
ARC
—
LiveB
—
Ctx
—
xAI · 4 months ago
Elo
1478
AA
62
GPQA
81.9
SWE
64
Term
52
ARC
20
LiveB
74
Ctx
256K
Fresh drops
Latest models.
Newest first — including rumored and announced models that aren't yet on the leaderboard.
- Sakana: Fugu Ultra v2· SakanaReleasedreleased · 6d ago
- DeepSeek: DeepSeek V4.1 Flash· DeepSeek
Hacker News · reports internal beta testing with native multimodal support.
Previewpreview · 7d ago - OpenAI: GPT-6 Astra· OpenAI
OpenAI · 2026-09-01 — Path to Astra: critical capabilities and frontier safeguards
Announcedannounced · 13d ago - OpenAI: GPT-6 Astra Pro· OpenAIReleasedreleased · 13d ago
- Qwen: Qwen3.8 Max (0902)· AlibabaReleasedreleased · 14d ago
- Releasedreleased · 15d ago
- Meta: Muse Spark 1.3· MetaReleasedreleased · 15d ago
- Google: Gemini 3.8 Flash· Google
r/singularity · 2026-09-02 · Introducing Gemini 3.8 Flash
Releasedreleased · 15d ago - Anthropic: Claude Fable 5.1· AnthropicReleasedreleased · 16d ago
- Tencent: Hy4 preview· TencentReleasedreleased · 20d ago
- Qwen: Qwen3.8 Flash· AlibabaReleasedreleased · 22d ago
- Z.ai: GLM 5.3 Flash· Z.aiReleasedreleased · 22d ago
- Releasedreleased · 27d ago
- DeepSeek: DeepSeek V4 Flash Vision Exp· DeepSeek
OpenRouter · preview endpoint live
Previewpreview · 27d ago - Z.ai: GLM 5.3· Z.ai
r/LocalLLaMA · official Z.ai announcement on 2026-08-14
Releasedreleased · 1mo ago
Score Sources
Aggregated by GPT-5.6 TerraEvery score in the table above is aggregated from these public leaderboards. Click any column header to sort the leaderboard; the ↗ icon opens the underlying source. Model names link to their OpenRouter page.
- LMArena
Chatbot Arena Elo · human head-to-head votes
- Artificial Analysis
Intelligence Index · aggregate of 10+ evals
- LiveBench
Contamination-free monthly benchmark
- GPQA Diamond
Grad-level science reasoning (198 Qs)
- SWE-bench Verified
Real GitHub issue fixes · human-audited
- Terminal-Bench
Agentic terminal tasks · Stanford / LAION
- ARC-AGI-2
Abstract reasoning · François Chollet
- Aider Polyglot
Multi-language code editing benchmark
- MMLU-Pro
Expanded multitask academic knowledge
- OpenRouter
Live model catalog · pricing · throughput
How to read this board
EditorialThe table above is a composite, not a poll. Each model's position comes from reconciling independent public evaluations — human preference votes on LMArena, the independently-run Artificial Analysis index, contamination-resistant LiveBench, graduate-level GPQA Diamond, and execution-graded coding benchmarks like SWE-bench Verified and Terminal-Bench. A model earns the top slot only when it leads across the widest set of those signals at once; winning a single benchmark is never enough.
Treat gaps of a few points as ties. Every evaluation on this page has a noise band of several points, and lab-reported figures regularly differ from independently-run ones — when they disagree, we weight the independent run. Below the top handful of models, scores are best read as estimates with the source attached, not as precise measurements.
"SOTA" is also time-stamped. The current leader, Anthropic: Claude Fable 5.1, has held the position for roughly 16 days since its public release. The frontier typically turns over every few weeks, so the date under the hero matters as much as the name: a SOTA claim from last month is a historical fact, not a current one.
If you are choosing a model rather than watching the race, don't start here — start with the job. The use-case leaderboards for coding, reasoning and agentic work re-weight the same underlying scores for each workload, and the methodology page documents every weight and guardrail.
FAQ
Common questions.
- What is the SOTA LLM right now?
- The current state-of-the-art LLM on our aggregated leaderboard is shown in the hero at the top of this page. Rankings are refreshed every couple of hours from LMArena, Artificial Analysis, LiveBench, GPQA, SWE-bench, Terminal-Bench and more.
- What does SOTA mean in AI?
- SOTA stands for state of the art — the model that currently leads on the benchmarks a task cares about. For a general SOTA we aggregate across every major public evaluation; for task-specific SOTA see /coding, /reasoning, or /agentic.
- Which LLM is best for coding?
- The dedicated coding leaderboard at /coding ranks models by a weighted composite of SWE-bench Verified, Aider Polyglot, Terminal-Bench and LiveBench — the benchmarks that actually predict developer productivity.
- How often is this leaderboard updated?
- Every ~2 hours. The exact last-refresh timestamp is shown under the H1 and in the footer of every page.
Signal Stream
24 items- Hacker News· 23m ago
Email marketing from Grok Bot official integration
Article URL: https://twitter.com/Migma_AI/status/2100617616704295197 Comments URL: https://news.ycombinator.com/item?id=49748310 Points: 1 # Comments: 0
- Simon Willison· 27m ago
How To Write With An LLM
How To Write With An LLM Thomas Ptacek on using LLMs as copyeditors, not as writing assistants: Rule Number One: You may not use a single word an LLM suggests to you. [...] I think that as a form of intellectual personal protective equipment you should adopt the rule that any specific turn of phrase an LLM suggests is off limits. Be strict about the rule! I won't let LLMs write content for my blog
- TechCrunch — AI· 38m ago
Crusoe raises $3.9B to build massive data centers and small modular “AI factories”
The round values the data center giant at $30.9 billion.
- TechCrunch — AI· 43m ago
Google DeepMind launches institute to widen the AGI debate
Google DeepMind just launched an institute to hash out the big AGI questions in public
- Hacker News· 58m ago
Alibaba releases Qwen 3.8 Omni Flash
Article URL: https://qwen.ai/blog?id=qwen3.8-omni-flash Comments URL: https://news.ycombinator.com/item?id=49747925 Points: 1 # Comments: 0
- TechCrunch — AI· 1h ago
PrismML hopes its tiny LLM will change how we all use AI
If AI lab PrismML isn't on your radar yet, it should be.
- TechCrunch — AI· 1h ago
The FAA’s plan to fix air traffic? $875M worth of AI
A new AI-based software program is being launched to help air traffic controllers better navigate their jobs as the crossing guards of America's skies.
- Ars Technica — AI· 1h ago
Small AI models let drones autonomously identify and attack battlefield targets
Scaleout deploys decentralized AI-driven learning to military bases and drones.
- r/singularity· 2h ago
How a small Israeli startup was linked to rogue AI hacks at OpenAI, Anthropic and Meta
Interesting detail behind the recent stories about OpenAI, Anthropic and Meta models “going rogue” during cyber evaluations. All three incidents were linked to the same third-party testing company, Irregular. According to Irregular, they resulted from the same evaluation-environment/containment issue — not a sophisticated sandbox escape.   submitted by   /u/AMBNNJ [link]   [comments]
- Hacker News· 2h ago
Show HN: Jev routing coding tasks to Grok Build or Codex Astra
Article URL: https://github.com/jcpsimmons/jev-model-router-demo Comments URL: https://news.ycombinator.com/item?id=49747155 Points: 2 # Comments: 0
- Hacker News· 2h ago
How Claude is uplifting biomolecular modeling
Article URL: https://www.anthropic.com/research/claude-uplifts-biomolecular-modeling Comments URL: https://news.ycombinator.com/item?id=49747128 Points: 2 # Comments: 0
- Hacker News· 2h ago
How to Write with an LLM
Article URL: https://sockpuppet.org/blog/2026/09/17/how-to-write-with-an-llm/ Comments URL: https://news.ycombinator.com/item?id=49747070 Points: 50 # Comments: 36
- r/singularity· 2h ago
Anthropic open-sources Claude-written GPU optimizations that make 30+ biomolecular models ~4× faster on average
  submitted by   /u/ResultBackground2450 [link]   [comments]
- r/singularity· 2h ago
What are others thinking and feeling about all this?
https://ai-2027.com/ https://www.reddit.com/r/singularity/s/Gmdb0LGfbR   submitted by   /u/Business_Way_8592 [link]   [comments]
- Simon Willison· 3h ago
Self-generated prompt injections in compaction summaries
Self-generated prompt injections in compaction summaries In Our framework for reporting model misalignment OpenAI provide "six reports on unexpected or concerning model behavior we’ve observed in the last six months". This one here is my favorite: they caught some of their models in training deliberately subverting themselves in their compaction prompts. Compaction is the process agent systems use
- r/singularity· 3h ago
Meet REBCO, the High-Temperature Superconducting Tape making Fusion possible, suitable for Magnets at 50 Teslas with a current density of 7 million amps per square centimeter -- and it relies on rare earth metals.
Oil? Pshah. Rare earths drive AI, solar cells, robotics, and now fusion.   submitted by   /u/Anen-o-me [link]   [comments]
- Hacker News· 3h ago
Using Jev for Claude Code model routing
Article URL: https://github.com/gargpratyush/jev-router Comments URL: https://news.ycombinator.com/item?id=49746321 Points: 2 # Comments: 0
- r/singularity· 3h ago
Anthropic reveals Claude is now leading 26% of its own R&D work, up from nearly zero 6 months ago
Blog: Measurements for understanding the pace of AI development inside frontier labs \ Anthropic   submitted by   /u/Outside-Iron-8242 [link]   [comments]
- Hacker News· 3h ago
Claude Slides, Claude Design and Claude Docs
Article URL: https://www.youtube.com/watch?v=To5nrYqvR44 Comments URL: https://news.ycombinator.com/item?id=49746238 Points: 2 # Comments: 0
- r/singularity· 3h ago
Measurements for understanding the pace of AI development inside frontier labs
  submitted by   /u/ResultBackground2450 [link]   [comments]
- TechCrunch — AI· 3h ago
The fix for rogue AI agents could be more AI
As companies hand off longer and more complex tasks to AI agents, they are running into an oversight problem: Agents can act faster, longer, and at greater volume than humans can realistically review.
- TechCrunch — AI· 3h ago
OpenAI caught its models leaving notes to successors to hide bad behavior
OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.
- Hacker News· 3h ago
Show HN: Vim-like, hyper-efficient, LLM power tool - 20-80k tokens not 350k+
I built a focused TUI tool for efficient parallelized tasks, early independent testers expressed 90%+ API savings and more direct results. This is accomplished by user smaller context windows, hybrid client-side compaction, tool compaction, and system nudges (to avoid polluting the base prompts). My Shell features: - Vim-inspired keybindings, your keys never have to leave the keyboard (some mouse
- Ars Technica — AI· 3h ago
Google announces new experimental "CC" AI agent for families
Multiple family members can share data to help the agent make plans and complete tasks.