这个评测关注什么
面向多步骤、长周期知识工作的专有评测。
阅读时注意
评测结果受任务集合、评分方法和模型版本影响。本站保留原始口径,不用它生成新的综合排名。
适用场景
用于回答一个明确的问题:模型在这一类任务上的相对表现如何,而不是“哪个模型绝对最好”。
长期知识工作
面向多步骤、长周期知识工作的专有评测。
| 排名 | 机构 | 模型 | Elo | 95% CI | 发布日期 |
|---|---|---|---|---|---|
| 1 | Anthropic | Claude Fable 5 | 1,583 | -15 / +16 | Jun 2026 |
| 2 | Kimi | Kimi K3 | 1,543 | -11 / +12 | Jul 2026 |
| 3 | OpenAI | GPT-5.6 Sol (max) | 1,496 | -12 / +12 | Jul 2026 |
| 4 | Anthropic | Claude Sonnet 5 (Adaptive Reasoning, Max Effort) | 1,388 | -12 / +11 | Jun 2026 |
| 5 | Anthropic | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | 1,354 | -11 / +10 | May 2026 |
| 6 | SpaceXAI | Grok 4.5 (high) | 1,323 | -12 / +13 | Jul 2026 |
| 7 | Anthropic | Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | 1,298 | -11 / +12 | Jun 2026 |
| 8 | Anthropic | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | 1,285 | -10 / +10 | Apr 2026 |
| 9 | Z AI | GLM-5.2 (max) | 1,260 | -10 / +11 | Jun 2026 |
| 10 | Anthropic | Claude Sonnet 5 (Adaptive Reasoning, High Effort) | 1,199 | -11 / +11 | Jun 2026 |
| 11 | OpenAI | GPT-5.5 (xhigh) | 1,154 | -9 / +8 | Apr 2026 |
| 12 | MiniMax | MiniMax-M3 | 1,110 | -9 / +10 | Jun 2026 |
| 13 | OpenAI | GPT-5.5 (high) | 1,103 | -9 / +9 | Apr 2026 |
| 14 | Anthropic | Claude Opus 4.7 (Non-reasoning, High Effort) | 1,089 | -10 / +10 | Apr 2026 |
| 15 | Anthropic | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) | 1,078 | -9 / +9 | Feb 2026 |
| 16 | Anthropic | Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) | 1,059 | -11 / +11 | Jun 2026 |
| 17 | OpenAI | GPT-5.5 (medium) | 1,000 | -0 / +0 | Apr 2026 |
| 18 | Z AI | GLM-5.1 (Reasoning) | 973 | -10 / +10 | Apr 2026 |
| 19 | DeepSeek | DeepSeek V4 Pro (Reasoning, Max Effort) | 932 | -10 / +9 | Apr 2026 |
| 20 | Anthropic | Claude Sonnet 5 (Adaptive Reasoning, Low Effort) | 926 | -10 / +10 | Jun 2026 |
| 21 | Alibaba | Qwen3.7 Max | 908 | -9 / +10 | May 2026 |
| 22 | Xiaomi | MiMo-V2.5-Pro | 873 | -10 / +9 | Apr 2026 |
| 23 | NVIDIA | Nemotron 3 Ultra 550B A55B (Reasoning) | 870 | -10 / +10 | Jun 2026 |
| 24 | Gemini 3.5 Flash (medium) | 869 | -9 / +10 | May 2026 | |
| 25 | Gemini 3.5 Flash (high) | 866 | -10 / +9 | May 2026 | |
| 26 | OpenAI | GPT-5.3 Codex (xhigh) | 864 | -10 / +10 | Feb 2026 |
| 27 | Meta | Muse Spark 1.1 (xhigh) | 863 | -13 / +12 | Jul 2026 |
| 28 | Thinking Machines | Inkling (xhigh) | 836 | -11 / +10 | Jul 2026 |
| 29 | DeepSeek | DeepSeek V4 Flash (Reasoning, Max Effort) | 831 | -9 / +9 | Apr 2026 |
| 30 | Kimi | Kimi K2.6 | 816 | -9 / +9 | Apr 2026 |
| 31 | Alibaba | Qwen3.6 27B (Reasoning) | 806 | -10 / +10 | Apr 2026 |
| 32 | SpaceXAI | Grok 4.3 (high) | 752 | -10 / +9 | Apr 2026 |
| 33 | OpenAI | GPT-5.4 mini (xhigh) | 707 | -11 / +9 | Mar 2026 |
| 34 | Meta | Muse Spark | 631 | -10 / +11 | Apr 2026 |
| 35 | Anthropic | Claude 4.5 Haiku (Reasoning) | 603 | -11 / +11 | Oct 2025 |
| 36 | KwaiKAT | KAT-Coder-Pro V1 | 592 | -12 / +12 | Nov 2025 |
| 37 | Alibaba | Qwen3.5 397B A17B (Reasoning) | 546 | -11 / +11 | Feb 2026 |
| 38 | Mistral | Mistral Medium 3.5 | 506 | -12 / +11 | Apr 2026 |
| 39 | Gemini 3.1 Pro Preview | 458 | -12 / +11 | Feb 2026 | |
| 40 | Gemma 4 31B (Reasoning) | 364 | -14 / +13 | Apr 2026 | |
| 41 | Cohere | Command A+ | 358 | -16 / +14 | May 2026 |
| 42 | Cohere | North Mini Code | 241 | -15 / +14 | Jun 2026 |
| 43 | Gemini 3.1 Flash-Lite | 221 | -15 / +14 | Mar 2026 | |
| 44 | Upstage | Solar Pro 3 | 128 | -15 / +14 | Apr 2026 |
| 45 | MBZUAI Institute of Foundation Models | K2 Think V2 | 50 | -16 / +14 | Dec 2025 |
| 46 | OpenAI | gpt-oss-120b (high) | 0 | -0 / +10 | Aug 2025 |
| 47 | OpenAI | gpt-oss-20b (high) | 0 | -0 / +0 | Aug 2025 |
| 48 | Meta | Llama 4 Maverick | 0 | -0 / +0 | Apr 2025 |
| 49 | NVIDIA | NVIDIA Nemotron 3 Super 120B A12B (Reasoning) | 0 | -0 / +0 | Mar 2026 |
面向多步骤、长周期知识工作的专有评测。
评测结果受任务集合、评分方法和模型版本影响。本站保留原始口径,不用它生成新的综合排名。
用于回答一个明确的问题:模型在这一类任务上的相对表现如何,而不是“哪个模型绝对最好”。