SOTA Model — the live state-of-the-art AI leaderboard. Current SOTA: Anthropic: Claude Fable 5 by Anthropic.
Current state of the art · v4.0 · streaming
Claude Fable 5 leads the aggregate across GPQA Diamond, SWE-bench Verified, Terminal-Bench, LiveBench, Aider Polyglot, MMLU-Pro, ARC-AGI-2, and Artificial Analysis Intelligence Index.
SOTA Elo
+14 vs #2 in the field
SOTA GPQA
Graduate-level science QA
Max Context
Token window at the top
Avg Output $
Across tracked frontier tier
Model Telemetry
| 01 | Anthropic: Claude Fable 5Anthropic SOTA | 1572 | 72 | 84.6 | 91.2 | 82.1 | 59.8 | 23.8 | 88.4 | 1M | $50.00 | 1 month ago |
| 02 | MoonshotAI: Kimi K3Moonshot | 1558 | 71 | 83.8 | 90.4 | 80.3 | 57.2 | 22.7 | 86.9 | 1.0M | $15.00 | 4 days ago |
| 03 | 1564 | 70 | 82.9 | 89.7 | 78.8 | 56.1 | 22.1 | 85.6 | 500K | $6.00 | 12 days ago | |
| 04 | DeepSeek: DeepSeek V4 ProDeepSeek | 1539 | 69 | 81.9 | 88.9 | 77.4 | 55.5 | 20.8 | 84.8 | 1.0M | $0.87 | 2 months ago |
| 05 | Qwen: Qwen3.7 MaxAlibaba | 1526 | 68 | 80.8 | 88.1 | 75.8 | 53.9 | 20.2 | 83.7 | 1M | $4.42 | 2 months ago |
| 06 | Anthropic: Claude Opus 4.8Anthropic | 1548 | 67 | 80.1 | 87.6 | 74.9 | 52.7 | 19.6 | 82.8 | 1M | $25.00 | 1 month ago |
| 07 | Anthropic: Claude Opus 4.8 (Fast)Anthropic | 1548 | 67 | 80.1 | 87.6 | 74.9 | 52.7 | 19.6 | 82.8 | 1M | $50.00 | 1 month ago |
| 08 | Z.ai: GLM 5.2Z.ai | 1518 | 66 | 79.4 | 86.9 | 73.6 | 51.8 | 18.9 | 81.9 | 1.0M | $2.98 | 1 month ago |
| 09 | NVIDIA: Nemotron 3 UltraNVIDIA | 1509 | 65 | 78.6 | 86.1 | 72.2 | 50.6 | 18.2 | 80.8 | 1M | $3.60 | 1 month ago |
| 10 | Sakana: Fugu UltraSakana | 1496 | 64 | 77.9 | 85.4 | 70.8 | 49.7 | 17.7 | 79.9 | 1M | $30.00 | 26 days ago |
| 11 | OpenAI: GPT-5.6 TerraOpenAI | 1502 | 63 | 77.2 | 84.8 | 69.9 | 48.9 | 17.1 | 79.1 | 1.1M | $15.00 | 11 days ago |
| 12 | 1505 | 63 | 77.5 | 85.1 | 70.4 | 49.3 | 17.3 | 79.4 | 1.1M | $15.00 | 11 days ago | |
| 13 | 1490 | 62 | 76.5 | 84.2 | 68.8 | 47.9 | 16.8 | 78.6 | 1.0M | $4.25 | 4 days ago | |
| 14 | MiniMax: MiniMax M3MiniMax | 1479 | 61 | 75.8 | 83.6 | 67.6 | 46.8 | 16.2 | 77.8 | 1.0M | $1.20 | 1 month ago |
| 15 | MoonshotAI: Kimi K2.7 CodeMoonshot | 1485 | 60 | 75.3 | 82.9 | 70.1 | 47.4 | 15.8 | 79.6 | 262K | $3.75 | 1 month ago |
| 16 | Z.ai: GLM 5.1Z.ai | 1468 | 59 | 74.5 | 82.1 | 66.9 | 45.9 | 15.3 | 77.1 | 203K | $3.04 | 3 months ago |
| 17 | 1473 | 58 | 73.9 | 81.7 | 65.8 | 45.2 | 14.9 | 76.4 | 1M | $2.50 | 2 months ago | |
| 18 | MoonshotAI: Kimi K2.6Moonshot | 1462 | 57 | 73.2 | 80.9 | 65.2 | 44.6 | 14.5 | 75.9 | 262K | $3.42 | 3 months ago |
| 19 | Mistral: Mistral Medium 3.5Mistral | 1451 | 56 | 72.3 | 80.1 | 63.7 | 43.4 | 13.8 | 74.8 | 262K | $7.50 | 2 months ago |
| 20 | Qwen: Qwen3.7 PlusAlibaba | 1443 | 55 | 71.6 | 79.4 | 62.6 | 42.7 | 13.4 | 73.9 | 1M | $1.28 | 1 month ago |
Anthropic · 1 month ago
Elo
1572
AA
72
GPQA
91.2
SWE
82
Term
60
ARC
24
LiveB
85
Ctx
1M
Moonshot · 4 days ago
Elo
1558
AA
71
GPQA
90.4
SWE
80
Term
57
ARC
23
LiveB
84
Ctx
1.0M
xAI · 12 days ago
Elo
1564
AA
70
GPQA
89.7
SWE
79
Term
56
ARC
22
LiveB
83
Ctx
500K
DeepSeek · 2 months ago
Elo
1539
AA
69
GPQA
88.9
SWE
77
Term
56
ARC
21
LiveB
82
Ctx
1.0M
Alibaba · 2 months ago
Elo
1526
AA
68
GPQA
88.1
SWE
76
Term
54
ARC
20
LiveB
81
Ctx
1M
Anthropic · 1 month ago
Elo
1548
AA
67
GPQA
87.6
SWE
75
Term
53
ARC
20
LiveB
80
Ctx
1M
Anthropic · 1 month ago
Elo
1548
AA
67
GPQA
87.6
SWE
75
Term
53
ARC
20
LiveB
80
Ctx
1M
Z.ai · 1 month ago
Elo
1518
AA
66
GPQA
86.9
SWE
74
Term
52
ARC
19
LiveB
79
Ctx
1.0M
NVIDIA · 1 month ago
Elo
1509
AA
65
GPQA
86.1
SWE
72
Term
51
ARC
18
LiveB
79
Ctx
1M
Sakana · 26 days ago
Elo
1496
AA
64
GPQA
85.4
SWE
71
Term
50
ARC
18
LiveB
78
Ctx
1M
OpenAI · 11 days ago
Elo
1502
AA
63
GPQA
84.8
SWE
70
Term
49
ARC
17
LiveB
77
Ctx
1.1M
OpenAI · 11 days ago
Elo
1505
AA
63
GPQA
85.1
SWE
70
Term
49
ARC
17
LiveB
78
Ctx
1.1M
Meta · 4 days ago
Elo
1490
AA
62
GPQA
84.2
SWE
69
Term
48
ARC
17
LiveB
77
Ctx
1.0M
MiniMax · 1 month ago
Elo
1479
AA
61
GPQA
83.6
SWE
68
Term
47
ARC
16
LiveB
76
Ctx
1.0M
Moonshot · 1 month ago
Elo
1485
AA
60
GPQA
82.9
SWE
70
Term
47
ARC
16
LiveB
75
Ctx
262K
Z.ai · 3 months ago
Elo
1468
AA
59
GPQA
82.1
SWE
67
Term
46
ARC
15
LiveB
75
Ctx
203K
xAI · 2 months ago
Elo
1473
AA
58
GPQA
81.7
SWE
66
Term
45
ARC
15
LiveB
74
Ctx
1M
Moonshot · 3 months ago
Elo
1462
AA
57
GPQA
80.9
SWE
65
Term
45
ARC
15
LiveB
73
Ctx
262K
Mistral · 2 months ago
Elo
1451
AA
56
GPQA
80.1
SWE
64
Term
43
ARC
14
LiveB
72
Ctx
262K
Alibaba · 1 month ago
Elo
1443
AA
55
GPQA
79.4
SWE
63
Term
43
ARC
13
LiveB
72
Ctx
1M
Score Sources
Aggregated by GPT-5.6 TerraEvery score in the table above is aggregated from these public leaderboards. Click any column header to sort the leaderboard; the ↗ icon opens the underlying source. Model names link to their OpenRouter page.
- LMArena
Chatbot Arena Elo · human head-to-head votes
- Artificial Analysis
Intelligence Index · aggregate of 10+ evals
- LiveBench
Contamination-free monthly benchmark
- GPQA Diamond
Grad-level science reasoning (198 Qs)
- SWE-bench Verified
Real GitHub issue fixes · human-audited
- Terminal-Bench
Agentic terminal tasks · Stanford / LAION
- ARC-AGI-2
Abstract reasoning · François Chollet
- Aider Polyglot
Multi-language code editing benchmark
- MMLU-Pro
Expanded multitask academic knowledge
- OpenRouter
Live model catalog · pricing · throughput
Signal Stream
24 items- Hacker News· 2h ago
Chrome-agent: LLM-native browser automation tool written in Rust
Article URL: https://github.com/sderosiaux/chrome-agent Comments URL: https://news.ycombinator.com/item?id=48982835 Points: 2 # Comments: 0
- Hacker News· 2h ago
Apply for Anthropic's AI for Science rare disease research grants
Article URL: https://www.anthropic.com/news/rare-disease-research-grants Comments URL: https://news.ycombinator.com/item?id=48981996 Points: 2 # Comments: 0
- Simon Willison· 3h ago
Who’s Afraid of Chinese Models?
Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and could help US open models compete more effectively with their Chinese counterparts: The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars
- Hacker News· 3h ago
Show HN: Amnesia – audit Claude Code's memory for contradictions
Article URL: https://github.com/tiny-cloud-ventures/amnesia Comments URL: https://news.ycombinator.com/item?id=48981446 Points: 3 # Comments: 0
- Hacker News· 3h ago
Show HN: Effort Router: Intelligent /effort selection per Claude turn
Dynamic effort selection, while possible in some other tools like OpenCode, does not exist in Claude Code. Having dynamic effort set per conversation turn means: - Less tokens burned on easy queries (e.g. something low /effort can do goes to xhigh because it's the convo default) - You avoid mistakes: Prompts requiring high effort (like a refactor or security threat modeling) go to high or xhigh in
- Hacker News· 4h ago
Chrome installed a global Ctrl+G keyboard shortcut to launch Gemini
Article URL: https://mastodon.online/users/mwichary/statuses/116952836351215165 Comments URL: https://news.ycombinator.com/item?id=48980930 Points: 11 # Comments: 0
- Hacker News· 4h ago
Claude Plays Robotics
Article URL: https://www.anthropic.com/research/claude-plays-robotics Comments URL: https://news.ycombinator.com/item?id=48980787 Points: 3 # Comments: 0
- The Verge — AI· 4h ago
Adobe’s ‘natural look’ camera app embraces generative AI
Adobe's experimental camera app has taken an unexpected turn. After Project Indigo was launched last year to provide a "more natural (SLR-like) look" for iPhone photography, the Indigo camera app is now being updated with a suite of generative AI tools. And the change doesn't rely upon Adobe's own Firefly AI models. Adobe describes the […]
- Hacker News· 4h ago
Hugging Face Turned to Chinese LLM for help after US models blocked Blue Team
Article URL: https://www.thestack.technology/hugging-face-hacked-turned-to-chinese-llm-for-help-after-us-models-blocked-blue-team/ Comments URL: https://news.ycombinator.com/item?id=48980470 Points: 2 # Comments: 0
- Hacker News· 5h ago
Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
Article URL: https://www.emergingtrajectories.com/lh/frontier-lab-economics/ Comments URL: https://news.ycombinator.com/item?id=48980019 Points: 172 # Comments: 172
- Hacker News· 5h ago
Surgical DevOps – Prevent LLM context drift and regressions
Article URL: https://github.com/bonushora/surgical-dev-ops/blob/main/README_EN.md Comments URL: https://news.ycombinator.com/item?id=48979553 Points: 3 # Comments: 0
- Hacker News· 5h ago
Logsum compresses log files to summary that matters, then ask your LLM
Article URL: https://github.com/flouthoc/logsum Comments URL: https://news.ycombinator.com/item?id=48979370 Points: 2 # Comments: 0
- Hacker News· 6h ago
Ask HN: Claude's Unexpected Blind Test: Users = Data = Value?
Comments URL: https://news.ycombinator.com/item?id=48979017 Points: 2 # Comments: 0
- Hacker News· 6h ago
Show HN: A Pipeline for Making 10-minute AI Movies with Claude Code and Seedance
I've been working on this Claude Code pipeline for making 10-minute movies, gluing together Seedance (videogen), Nano Banana (image gen), and ElevenLabs (voice gen), and using Claude as the director. My repo contributes a markdown playbook and a worked example which should make it easier for other people to make their own short movies. A first-pass movie takes around 2.5 hours of wall clock time a
- Hacker News· 6h ago
Deterministic LLM router with a cryptographic audit trail – TEIA
Article URL: https://github.com/felippebarcelos/teia-omega-awakening Comments URL: https://news.ycombinator.com/item?id=48978938 Points: 2 # Comments: 0
- Hacker News· 6h ago
Using LLM-Based Verification to Eliminate Bugs in Linux's Network Stack
Article URL: https://www.basis.ai/blog/verified-nftables/ Comments URL: https://news.ycombinator.com/item?id=48978901 Points: 3 # Comments: 0
- Hacker News· 7h ago
AgentAbstain: Do LLM Agents Know When Not to Act?
Article URL: https://arxiv.org/abs/2607.10059 Comments URL: https://news.ycombinator.com/item?id=48978098 Points: 3 # Comments: 0
- Hacker News· 8h ago
Automating first-pass customer support with Claude Code and MCP
Article URL: https://sitespeak.ai/blog/automate-customer-support-mcp Comments URL: https://news.ycombinator.com/item?id=48977582 Points: 2 # Comments: 0
- The Verge — AI· 10h ago
China delivers a one-two punch to America’s AI dominance
China's leading AI companies are ramping up the pressure on Silicon Valley, as Moonshot and Alibaba unveiled models they claim can go toe-to-toe with the best from OpenAI and Anthropic at a fraction of the cost. The rapid-fire releases suggest America's lead at the AI frontier is increasingly tight, just as the technology is becoming […]
- OpenAI· 10h ago
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
- Hacker News· 10h ago
How to Write Loops with Claude Code
Article URL: https://karanbansal.in/blog/loop-engineering/ Comments URL: https://news.ycombinator.com/item?id=48976424 Points: 1 # Comments: 0
- Hacker News· 11h ago
Agentic test processes, LLM benchmarks, and other notes from Galapagos Island
Article URL: https://danluu.com/ai-coding/ Comments URL: https://news.ycombinator.com/item?id=48976284 Points: 1 # Comments: 0
- Hacker News· 11h ago
LLM Wiki Implementation
Article URL: https://github.com/nashsu/llm_wiki Comments URL: https://news.ycombinator.com/item?id=48976240 Points: 1 # Comments: 0
- Hacker News· 12h ago
Free GPT 5.6 Terra and Luna (2.5M/day) & 250k Sol – free experiments
Article URL: https://platform.openai.com/settings/organization/data-controls/sharing Comments URL: https://news.ycombinator.com/item?id=48975696 Points: 2 # Comments: 1