LLM 基准排行榜
来自 Wikiprompt,自由的提示词百科
大语言模型与生成式 AI 的实时排行榜,每日从开放的社区评测汇总。每张表都注明来源,Wikiprompt 每天保存快照,排名的变化都有记录。
最后更新: 2026-08-24 · 这是实时排行榜;带日期的快照保留每一次过去的排名。 历史 →
交互式浏览器
按竞技场、机构或许可证筛选,搜索模型,排序表格,并查看成本与性能图。下方的表格仍是可引用的参考。
文本竞技场
点击表格行或图上的点来固定模型,只对比选中的模型。
成本与性能
输出价格(美元/100万token,对数刻度,来自 OpenRouter)对比社区 Elo(LMArena)。左上角 = 性价比最高。
Elo 随时间变化(文本类前几名)
| 排名 ▲ | 模型 | 机构 | Elo 分数 | 投票数 | 输入 $/1M | 输出 $/1M | 上下文 |
|---|---|---|---|---|---|---|---|
| 1 | claude-opus-5-high | anthropic | 1,504 | 27,610 | $10 | $50 | 1,000,000 |
| 3 | claude-opus-4-6-high | anthropic | 1,503 | 72,359 | $5 | $25 | 1,000,000 |
| 4 | claude-opus-4-6 | anthropic | 1,497 | 76,362 | $5 | $25 | 1,000,000 |
| 5 | claude-fable-5 | anthropic | 1,495 | 24,331 | $10 | $50 | 1,000,000 |
| 7 | claude-opus-4-7-high | anthropic | 1,490 | 60,203 | $5 | $25 | 1,000,000 |
| 9 | claude-opus-4-7 | anthropic | 1,483 | 61,304 | $5 | $25 | 1,000,000 |
| 10 | gemini-3.5-flash-high | 1,483 | 30,004 | $1.5 | $9 | 1,048,576 | |
| 12 | gemini-3.1-pro-preview | 1,480 | 99,182 | $2 | $12 | 1,048,576 | |
| 13 | gemini-3-pro | 1,479 | 41,491 | $2 | $12 | 131,072 | |
| 14 | muse-spark-1.1 | meta | 1,478 | 20,242 | $1.25 | $4.25 | 1,048,576 |
| 18 | gemini-3.5-flash-medium | 1,475 | 28,343 | $1.5 | $9 | 1,048,576 | |
| 21 | qwen3.5-max-preview | alibaba | 1,471 | 21,484 | |||
| 22 | gpt-5.5-high | openai | 1,471 | 59,545 | $5 | $30 | 1,050,000 |
| 23 | gpt-5.4-high | openai | 1,470 | 60,708 | $2.5 | $15 | 1,050,000 |
| 24 | ernie-5.1 | baidu | 1,468 | 37,137 | |||
| 25 | gemini-3-flash | 1,466 | 30,852 | $0.5 | $3 | 1,048,576 | |
| 26 | gpt-5.5 | openai | 1,466 | 60,896 | $5 | $30 | 1,050,000 |
| 27 | glm-5.2-max | zai | 1,465 | 30,637 | $0.97 | $3.04 | 1,048,576 |
| 28 | mimo-v2.5-pro | xiaomi | 1,465 | 54,177 | $0.44 | $0.87 | 1,050,000 |
| 29 | glm-5.1 | zai | 1,464 | 42,267 | $0.97 | $3.04 | 204,800 |
| 30 | claude-opus-4-8-high | anthropic | 1,461 | 44,676 | $5 | $25 | 1,000,000 |
| 31 | claude-sonnet-4-6 | anthropic | 1,458 | 66,512 | $2 | $10 | 1,000,000 |
| 32 | gemini-2.5-pro | 1,457 | 124,807 | $1.25 | $10 | 1,048,576 | |
| 33 | qwen3.7-plus | alibaba | 1,456 | 34,426 | $0.32 | $1.28 | 1,000,000 |
| 34 | kimi-k2.6 | moonshot | 1,455 | 37,543 | $0.95 | $4 | 262,144 |
| 35 | gpt-5.6-sol-xhigh | openai | 1,454 | 19,541 | $2 | $10 | 1,050,000 |
| 36 | gpt-5.4 | openai | 1,453 | 63,659 | $2.5 | $15 | 1,050,000 |
| 37 | grok-4.5 | xai | 1,452 | 22,030 | $2 | $6 | 500,000 |
| 38 | claude-opus-4-8 | anthropic | 1,452 | 45,331 | $5 | $25 | 1,000,000 |
| 39 | grok-4.20-beta-0309-reasoning | xai | 1,451 | 62,272 | $1.25 | $2.5 | 2,000,000 |
| 40 | deepseek-v4-pro | deepseek | 1,451 | 54,263 | $0.52 | $1.04 | 1,048,576 |
| 41 | grok-4.20-multi-agent-beta-0309 | xai | 1,450 | 60,864 | $1.25 | $2.5 | 2,000,000 |
| 42 | claude-opus-4-5-20251101 | anthropic | 1,450 | 70,985 | $5 | $25 | 1,000,000 |
| 43 | dola-seed-2.0-pro | bytedance | 1,448 | 74,513 | |||
| 45 | claude-opus-4-5-20251101-high-32k | anthropic | 1,447 | 37,198 | $5 | $25 | 1,000,000 |
| 47 | gpt-5.6-terra-xhigh | openai | 1,446 | 20,216 | $2 | $12 | 1,050,000 |
| 48 | glm-5 | zai | 1,445 | 27,850 | $0.6 | $1.92 | 204,800 |
| 49 | kimi-k2.5-thinking | moonshot | 1,445 | 70,943 | $0.45 | $2.25 | 262,144 |
| 50 | deepseek-v4-pro-high-preview | deepseek | 1,445 | 51,642 | $0.52 | $1.04 | 1,048,576 |
| 52 | ernie-5.0-0110 | baidu | 1,444 | 35,267 | |||
| 53 | grok-4.20-beta1 | xai | 1,444 | 26,770 | $1.25 | $2.5 | 2,000,000 |
| 56 | gemini-3-flash (thinking-minimal) | 1,442 | 86,359 | $0.5 | $3 | 1,048,576 | |
| 57 | claude-sonnet-5-high | anthropic | 1,442 | 27,366 | $2 | $10 | 1,000,000 |
| 59 | gpt-5.1-high | openai | 1,441 | 40,945 | $1.25 | $10 | 400,000 |
| 62 | glm-4.6 | zai | 1,440 | 35,830 | $0.5 | $2 | 204,800 |
| 65 | qwen3-max-preview | alibaba | 1,439 | 27,838 | $0.78 | $3.9 | 262,144 |
| 66 | gpt-5.2-chat-latest-20260210 | openai | 1,439 | 34,339 | $1.75 | $14 | 128,000 |
| 67 | qwen3.5-397b-a17b | alibaba | 1,438 | 70,207 | $0.5 | $3.6 | 262,144 |
| 68 | claude-sonnet-4-5-20250929 | anthropic | 1,438 | 80,858 | $2 | $10 | 1,000,000 |
| 69 | grok-4.1-thinking | xai | 1,437 | 65,550 | |||
| 70 | qwen3.6-plus | alibaba | 1,437 | 45,376 | $0.33 | $1.95 | 1,000,000 |
| 72 | mimo-v2-pro | xiaomi | 1,437 | 24,397 | |||
| 74 | grok-4.1 | xai | 1,436 | 67,767 | |||
| 76 | minimax-m3 | minimax | 1,435 | 41,023 | $0.3 | $1.2 | 1,048,576 |
| 78 | claude-sonnet-4-5-20250929-high-32k | anthropic | 1,433 | 82,473 | $2 | $10 | 1,000,000 |
| 79 | deepseek-v4-flash | deepseek | 1,432 | 49,001 | $0.04 | $0.08 | 1,310,720 |
| 80 | glm-4.5 | zai | 1,429 | 24,403 | $0.6 | $2.2 | 131,072 |
| 82 | chatgpt-4o-latest-20250326 | openai | 1,429 | 82,766 | |||
| 83 | gpt-5.6-luna-xhigh | openai | 1,429 | 20,693 | $0.2 | $1.2 | 1,050,000 |
| 84 | mistral-large-3 | mistral | 1,428 | 62,644 | $0.5 | $1.5 | 262,144 |
价格按归一化模型名与最接近的 OpenRouter 条目匹配;推理强度变体共享基础模型价格。
要点
- • 当前文本竞技场领先者: claude-opus-5-high (anthropic, Elo 1504)
- • 最佳性价比(每美元Elo,输出低于$2/1M): mimo-v2.5-pro (Elo 1465, $0.87/1M)
- • 最强开源模型: glm-5.2-max (MIT, #27)
- • 编程领先者(Aider通过率): gpt-5 (high) (88%)
文本竞技场25 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | claude-opus-5-high | anthropic | 1,504 |
| 3 | claude-opus-4-6-high | anthropic | 1,503 |
| 4 | claude-opus-4-6 | anthropic | 1,497 |
| 5 | claude-fable-5 | anthropic | 1,495 |
| 7 | claude-opus-4-7-high | anthropic | 1,490 |
| 9 | claude-opus-4-7 | anthropic | 1,483 |
| 10 | gemini-3.5-flash-high | 1,483 | |
| 12 | gemini-3.1-pro-preview | 1,480 | |
| 13 | gemini-3-pro | 1,479 | |
| 14 | muse-spark-1.1 | meta | 1,478 |
| 18 | gemini-3.5-flash-medium | 1,475 | |
| 21 | qwen3.5-max-preview | alibaba | 1,471 |
| 22 | gpt-5.5-high | openai | 1,471 |
| 23 | gpt-5.4-high | openai | 1,470 |
| 24 | ernie-5.1 | baidu | 1,468 |
| 25 | gemini-3-flash | 1,466 | |
| 26 | gpt-5.5 | openai | 1,466 |
| 27 | glm-5.2-max | zai | 1,465 |
| 28 | mimo-v2.5-pro | xiaomi | 1,465 |
| 29 | glm-5.1 | zai | 1,464 |
| 30 | claude-opus-4-8-high | anthropic | 1,461 |
| 31 | claude-sonnet-4-6 | anthropic | 1,458 |
| 32 | gemini-2.5-pro | 1,457 | |
| 33 | qwen3.7-plus | alibaba | 1,456 |
| 34 | kimi-k2.6 | moonshot | 1,455 |
Web 开发竞技场25 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | claude-opus-5-max | anthropic | 1,691 |
| 2 | kimi-k3-max | moonshot | 1,674 |
| 3 | qwen3.8-max | alibaba | 1,669 |
| 4 | claude-opus-5-high | anthropic | 1,663 |
| 5 | grok-4.6-high | xai | 1,629 |
| 6 | claude-fable-5 | anthropic | 1,626 |
| 7 | gpt-5.6-sol-xhigh (codex-harness) | openai | 1,619 |
| 8 | glm-5.3-max | zai | 1,599 |
| 9 | qwen3.8-27b | alibaba | 1,595 |
| 10 | gemini-3.7-flash-high | 1,587 | |
| 11 | glm-5.2-max | zai | 1,582 |
| 12 | deepseek-v4-pro-high-20260813 | deepseek | 1,582 |
| 13 | deepseek-v4-flash-high | deepseek | 1,579 |
| 14 | claude-opus-4-8-high | anthropic | 1,563 |
| 15 | claude-opus-4-7 | anthropic | 1,558 |
| 16 | claude-opus-4-7-high | anthropic | 1,557 |
| 17 | grok-4.5 | xai | 1,556 |
| 18 | claude-opus-4-6-high | anthropic | 1,546 |
| 19 | claude-opus-4-8 | anthropic | 1,539 |
| 20 | muse-spark-1.1 | meta | 1,539 |
| 21 | gemini-3.6-flash-high | 1,539 | |
| 22 | claude-sonnet-5-high | anthropic | 1,539 |
| 23 | claude-opus-4-6 | anthropic | 1,536 |
| 24 | muse-spark-1.2 (xHigh) | meta | 1,534 |
| 25 | claude-sonnet-4-6 | anthropic | 1,522 |
视觉竞技场25 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | claude-fable-5 | anthropic | 1,328 |
| 2 | claude-opus-5-high | anthropic | 1,323 |
| 3 | claude-opus-4-7 | anthropic | 1,317 |
| 4 | claude-opus-4-7-high | anthropic | 1,316 |
| 5 | qwen3.8-max | alibaba | 1,315 |
| 6 | claude-opus-4-6-high | anthropic | 1,315 |
| 7 | gemini-3.5-flash-high | 1,313 | |
| 8 | claude-opus-4-6 | anthropic | 1,311 |
| 9 | gemini-3.5-flash-medium | 1,308 | |
| 10 | muse-spark | meta | 1,306 |
| 11 | gemini-3-pro | 1,305 | |
| 12 | muse-spark-1.2 (xHigh) | meta | 1,304 |
| 13 | gemini-3.6-flash-high | 1,302 | |
| 14 | gpt-5.4-high | openai | 1,297 |
| 15 | gpt-5.5 | openai | 1,295 |
| 16 | muse-spark-1.1 | meta | 1,295 |
| 17 | gemini-3.1-pro-preview | 1,294 | |
| 18 | claude-opus-4-8-high | anthropic | 1,293 |
| 19 | gpt-5.5-high | openai | 1,293 |
| 20 | gpt-5.4 | openai | 1,293 |
| 21 | grok-4.5 | xai | 1,291 |
| 22 | claude-opus-4-8 | anthropic | 1,287 |
| 23 | gemini-3-flash | 1,283 | |
| 24 | claude-sonnet-4-6 | anthropic | 1,281 |
| 25 | claude-sonnet-5-high | anthropic | 1,281 |
文生图竞技场25 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | gpt-image-2 (medium) | openai | 1,380 |
| 2 | reve-2.1 | reve | 1,302 |
| 3 | muse-image | meta | 1,283 |
| 4 | reve-2.0 | reve | 1,270 |
| 5 | gemini-3.1-flash-image (nano-banana-2) [web-search] | 1,263 | |
| 6 | qwen-image-3.0-pro | alibaba | 1,258 |
| 7 | seedream-5.0-pro | bytedance | 1,257 |
| 8 | mai-image-2.5 | microsoft-ai | 1,256 |
| 9 | gemini-3.1-flash-lite-image (nano-banana-2-lite) | 1,251 | |
| 10 | gemini-3-pro-image-2k (nano-banana-pro) | 1,246 | |
| 11 | gpt-image-1.5-high-fidelity | openai | 1,239 |
| 12 | gemini-3-pro-image-preview (nano-banana-pro) | 1,232 | |
| 13 | grok-imagine-image-quality | xai | 1,228 |
| 14 | ideogram-4.0-quality | ideogram | 1,206 |
| 15 | qwen-image-2.0-pro-2026-06-22 | alibaba | 1,191 |
| 16 | uni-1.1-max | luma-ai | 1,188 |
| 17 | mai-image-2 | microsoft-ai | 1,182 |
| 18 | Cosmos3-Super-Text2Image (Agentic) | nvidia | 1,181 |
| 19 | uni-1.1 | luma-ai | 1,180 |
| 20 | grok-imagine-image | xai | 1,172 |
| 21 | recraft-v4.1-utility-pro | recraft | 1,169 |
| 22 | flux-2-max | bfl | 1,162 |
| 23 | grok-imagine-image-pro | xai | 1,161 |
| 24 | Cosmos3-Super-Text2Image | nvidia | 1,158 |
| 25 | flux-2-flex | bfl | 1,156 |
图像编辑竞技场25 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | gpt-image-2 (medium) | openai | 1,463 |
| 2 | grok-imagine-image-2.0 (low) | xai | 1,439 |
| 3 | mai-image-2.6-preview | microsoft-ai | 1,420 |
| 4 | muse-image | meta | 1,406 |
| 5 | mai-image-2.5 | microsoft-ai | 1,401 |
| 6 | seedream-5.0-pro | bytedance | 1,394 |
| 7 | grok-imagine-image-quality (20260519) | xai | 1,390 |
| 8 | gemini-3-pro-image-2k (nano-banana-pro) | 1,390 | |
| 9 | chatgpt-image-latest-high-fidelity (20251216) | openai | 1,389 |
| 10 | gemini-3.1-flash-image (nano-banana-2) [web-search] | 1,386 | |
| 11 | gemini-3-pro-image-preview (nano-banana-pro) | 1,385 | |
| 12 | reve-2.1 | reve | 1,375 |
| 13 | gpt-image-1.5-high-fidelity | openai | 1,370 |
| 14 | reve-2.0 | reve | 1,358 |
| 15 | uni-1.1-max | luma-ai | 1,334 |
| 16 | grok-imagine-image | xai | 1,330 |
| 17 | gemini-3.1-flash-lite-image (nano-banana-2-lite) | 1,314 | |
| 18 | uni-1.1 | luma-ai | 1,312 |
| 19 | qwen-image-2.0-pro-2026-06-22 | alibaba | 1,304 |
| 20 | hunyuan-image-3.0-instruct | tencent | 1,303 |
| 21 | wan2.7-image-pro | alibaba | 1,302 |
| 22 | seedream-4.5 | bytedance | 1,302 |
| 23 | wan2.7-image | alibaba | 1,301 |
| 24 | gemini-2.5-flash-image-preview (nano-banana) | 1,293 | |
| 25 | seedream-5.0-lite | bytedance | 1,293 |
文生视频竞技场25 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | gemini-omni-flash | 1,512 | |
| 2 | flux-3-video | bfl | 1,494 |
| 3 | dreamina-seedance-2.0-720p | bytedance | 1,482 |
| 4 | dreamina-seedance-2.5-720p | bytedance | 1,477 |
| 5 | muse-video | meta | 1,457 |
| 6 | minimax-h3 | minimax | 1,453 |
| 7 | happyhorse-1.0 | aorizon | 1,428 |
| 8 | sora-2-pro | openai | 1,364 |
| 9 | veo-3.1-audio | 1,364 | |
| 10 | veo-3.1-audio-1080p | 1,363 | |
| 11 | veo-3.1-fast-audio | 1,361 | |
| 12 | veo-3.1-fast-audio-1080p | 1,358 | |
| 13 | veo-3-fast-audio | 1,347 | |
| 14 | grok-imagine-video-720p | xai | 1,345 |
| 15 | wan2.7-t2v | wan | 1,344 |
| 16 | sora-2 | openai | 1,340 |
| 17 | veo-3-audio | 1,339 | |
| 18 | wan2.6-t2v | alibaba | 1,330 |
| 19 | seedance-v1.5-pro | bytedance | 1,256 |
| 20 | veo-3 | 1,253 | |
| 21 | wan2.5-t2v-preview | alibaba | 1,249 |
| 22 | veo-3-fast | 1,248 | |
| 23 | pixverse-v5.6 | 1,240 | |
| 24 | runway-gen-4.5 | runway | 1,223 |
| 25 | kling-2.5-turbo-1080p | kling | 1,219 |
视频编辑竞技场9 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | dreamina-seedance-2.5-720p | bytedance | 1,411 |
| 2 | minimax-h3 | minimax | 1,388 |
| 3 | dreamina-seedance-2.0-720p | bytedance | 1,359 |
| 4 | gemini-omni-flash | 1,358 | |
| 5 | happyhorse-1.0 | aorizon | 1,307 |
| 6 | grok-imagine-video | xai | 1,262 |
| 7 | kling-o3-pro | kling | 1,257 |
| 8 | kling-o1-pro | kling | 1,200 |
| 9 | runway-gen4-aleph | runway | 1,183 |
搜索竞技场25 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | claude-opus-4-6-search | anthropic | 1,253 |
| 2 | gpt-5.5-search | openai | 1,240 |
| 3 | claude-fable-5 | anthropic | 1,237 |
| 4 | claude-opus-4-7 | anthropic | 1,233 |
| 5 | ernie-5.1 | baidu | 1,226 |
| 6 | claude-sonnet-4-6-search | anthropic | 1,221 |
| 7 | gemini-3.1-pro-grounding | 1,212 | |
| 8 | gemini-3-pro-grounding | 1,207 | |
| 9 | gpt-5.2-search | openai | 1,206 |
| 10 | grok-4.20-multi-agent-beta-0309 | xai | 1,205 |
| 11 | claude-opus-4-8 | anthropic | 1,205 |
| 12 | gpt-5.1-search | openai | 1,199 |
| 13 | gemini-3-flash-grounding | 1,197 | |
| 14 | gpt-5.4-search | openai | 1,196 |
| 15 | grok-4.20-beta1 | xai | 1,189 |
| 16 | claude-sonnet-5-search | anthropic | 1,188 |
| 17 | claude-opus-4-5-search | anthropic | 1,179 |
| 18 | gpt-5.2-search-non-reasoning | openai | 1,173 |
| 19 | grok-4-1-fast-search | xai | 1,171 |
| 20 | grok-4-fast-search | xai | 1,170 |
| 21 | grok-4.3 | xai | 1,163 |
| 22 | claude-sonnet-4-5-search | anthropic | 1,158 |
| 23 | claude-opus-4-1-search | anthropic | 1,148 |
| 24 | o3-search | openai | 1,143 |
| 25 | gemini-2.5-pro-grounding | 1,142 |
智能体竞技场25 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | Claude Opus 5 (High) | anthropic | |
| 2 | Claude Opus 5 (Max) | anthropic | |
| 3 | Claude Fable 5 (High) | anthropic | |
| 4 | Kimi K3 (Max) | moonshot | |
| 5 | GPT 5.6 Sol (xHigh) | openai | |
| 6 | Claude Opus 4.8 (High) | anthropic | |
| 7 | GPT 5.5 (xHigh) | openai | |
| 8 | Claude Opus 4.7 (High) | anthropic | |
| 9 | GPT 5.5 (High) | openai | |
| 10 | Claude Opus 4.7 | anthropic | |
| 11 | Claude Sonnet 5 (High) | anthropic | |
| 12 | Claude Opus 4.6 | anthropic | |
| 13 | GPT 5.5 | openai | |
| 14 | DeepSeek V4 Pro (High) (0813) | deepseek | |
| 15 | Qwen3.8 Max | alibaba | |
| 16 | Grok 4.5 | xai | |
| 17 | GLM 5.2 (Max) | zai | |
| 18 | GPT 5.4 (High) | openai | |
| 19 | GPT 5.6 Luna (xHigh) | openai | |
| 20 | Deepseek V4 Flash (High) (20260731) | deepseek | |
| 21 | Gemini 3.7 Flash (High) | ||
| 22 | GPT 5.6 Terra (xHigh) | openai | |
| 23 | Claude Sonnet 4.6 | anthropic | |
| 24 | Claude Opus 4.8 | anthropic | |
| 25 | Muse Spark 1.1 | meta |
文档竞技场25 个模型完整排行榜 →
| 排名 | 模型 | 机构 | Elo 分数 |
|---|---|---|---|
| 1 | claude-opus-5-high | anthropic | 1,520 |
| 2 | claude-opus-4-6 | anthropic | 1,510 |
| 3 | claude-opus-4-6-thinking | anthropic | 1,506 |
| 4 | claude-fable-5 | anthropic | 1,504 |
| 5 | claude-opus-4-7 | anthropic | 1,498 |
| 6 | claude-opus-4-7-thinking | anthropic | 1,497 |
| 7 | gpt-5.5-high | openai | 1,485 |
| 8 | claude-sonnet-4-6 | anthropic | 1,483 |
| 9 | gpt-5.5 | openai | 1,480 |
| 10 | gpt-5.6-terra-xhigh | openai | 1,479 |
| 11 | gpt-5.6-sol-xhigh | openai | 1,479 |
| 12 | claude-opus-4-8-thinking | anthropic | 1,475 |
| 13 | muse-spark-1.1 | meta | 1,472 |
| 14 | claude-sonnet-5-high | anthropic | 1,470 |
| 15 | gpt-5.4 | openai | 1,470 |
| 16 | claude-opus-4-8 | anthropic | 1,469 |
| 17 | gemini-3.5-flash-medium | 1,465 | |
| 18 | gpt-5.6-luna-xhigh | openai | 1,462 |
| 19 | claude-opus-4-5-20251101 | anthropic | 1,462 |
| 20 | grok-4.5 | xai | 1,454 |
| 21 | kimi-k2.6 | moonshot | 1,451 |
| 22 | claude-sonnet-4-5-20250929 | anthropic | 1,446 |
| 23 | gemini-3.1-pro-preview | 1,445 | |
| 24 | muse-spark | meta | 1,443 |
| 25 | qwen3.7-plus | alibaba | 1,440 |
其他排行榜
LiveBench
指标: average score (0-100) · livebench.ai
| 排名 | 模型 | 机构 | 分数 |
|---|---|---|---|
| 1 | claude-fable-5-max-effort | 83.4 | |
| 2 | gpt-5.6-sol-max | 81.7 | |
| 3 | gpt-5.5-xhigh | 80.8 | |
| 4 | claude-opus-5-max-effort | 80.5 | |
| 5 | gemini-3.7-flash-high | 79.9 | |
| 6 | smaug-agentic | 79.7 | |
| 7 | qwen3.8-max | 79.5 | |
| 8 | kimi-k3 | 79.5 | |
| 9 | grok-4.6 | 79 | |
| 10 | muse-spark-1.2-xhigh | 78.9 | |
| 11 | gpt-5.4-xhigh | 78.8 | |
| 12 | gpt-5.6-terra-max | 78.6 | |
| 13 | deepseek-v4-pro-0813 | 78.2 | |
| 14 | gemini-3.1-pro-preview-high | 78 | |
| 15 | deepseek-v4-flash-vision-exp | 77.7 |
BenchLM
指标: aggregated display score · benchlm.ai
| 排名 | 模型 | 机构 | 分数 |
|---|---|---|---|
| 1 | Claude Mythos 5 | Anthropic | 83 |
| 2 | Claude Opus 5 | Anthropic | 82.7 |
| 3 | Claude Fable 5 | Anthropic | 82.7 |
| 4 | GPT-5.6 Sol | OpenAI | 81.7 |
| 5 | Kimi K3 | Moonshot AI | 80.2 |
| 6 | Qwen3.8 Max | Alibaba | 79 |
| 7 | Muse Spark 1.1 | Meta | 76.7 |
| 8 | Claude Opus 4.8 | Anthropic | 76.2 |
| 9 | Gemini 3.6 Flash | 75.2 | |
| 10 | Grok 4.5 | xAI | 75.2 |
| 11 | GPT-5.5 | OpenAI | 73.3 |
| 12 | GPT-5.4 | OpenAI | 73.1 |
| 13 | GPT-5.6 Terra | OpenAI | 72.6 |
| 14 | Claude Opus 4.7 (Adaptive) | Anthropic | 72.3 |
| 15 | Claude Opus 4.7 | Anthropic | 71.9 |
Aider polyglot (coding)
指标: pass rate 2 (%) · aider.chat/docs/leaderboards/
| 排名 | 模型 | 机构 | 分数 |
|---|---|---|---|
| 1 | gpt-5 (high) | 88 | |
| 2 | gpt-5 (medium) | 86.7 | |
| 3 | o3-pro (high) | 84.9 | |
| 4 | gemini-2.5-pro-preview-06-05 (32k think) | 83.1 | |
| 5 | o3 (high) | 81.3 | |
| 6 | gpt-5 (low) | 81.3 | |
| 7 | grok-4 (high) | 79.6 | |
| 8 | gemini-2.5-pro-preview-06-05 (default think) | 79.1 | |
| 9 | o3 (high) + gpt-4.1 | 78.2 | |
| 10 | Gemini 2.5 Pro Preview 05-06 | 76.9 | |
| 11 | o3 | 76.9 | |
| 12 | DeepSeek-V3.2-Exp (Reasoner) | 74.2 | |
| 13 | Gemini 2.5 Pro Preview 03-25 | 72.9 | |
| 14 | o4-mini (high) | 72 | |
| 15 | claude-opus-4-20250514 (32k thinking) | 72 |
Artificial Analysis (Intelligence Index)
指标: AA Intelligence Index (0-100) · artificialanalysis.ai
| 排名 | 模型 | 机构 | 分数 |
|---|---|---|---|
| 1 | Claude Opus 5 (Adaptive Reasoning, Max Effort) | Anthropic | 63.1 |
| 2 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 62.5 |
| 3 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 62.1 |
| 4 | Claude Opus 5 (Adaptive Reasoning, High Effort) | Anthropic | 61.5 |
| 5 | GPT-5.6 Sol (max) | OpenAI | 60.9 |
| 6 | Grok 4.6 (high) | SpaceXAI | 60.9 |
| 7 | Grok 4.6 (xhigh) | SpaceXAI | 60 |
| 8 | Kimi K3 (max) | Kimi | 59.7 |
| 9 | GLM-5.3 (max) | Z AI | 59.5 |
| 10 | GPT-5.6 Sol (xhigh) | OpenAI | 59 |
| 11 | Grok 4.6 (medium) | SpaceXAI | 59 |
| 12 | Claude Opus 5 (Adaptive Reasoning, Medium Effort) | Anthropic | 58.6 |
| 13 | Qwen3.8 Max | Alibaba | 58.1 |
| 14 | Qwen3.8 2.4T A95B | Alibaba | 57.7 |
| 15 | GPT-5.6 Sol (high) | OpenAI | 57.3 |
价格与上下文窗口
价格与上下文元数据来自 OpenRouter 公开 API。价格为每 100 万 token 的美元数。
| 模型 | 机构 | 上下文 | 输入 $/1M | 输出 $/1M |
|---|---|---|---|---|
| Auto Router (Beta) | openrouter | 2,000,000 | $-1,000,000 | $-1,000,000 |
| Pareto Code Router | openrouter | 2,000,000 | $-1,000,000 | $-1,000,000 |
| SpaceXAI: Grok 4.20 Multi-Agent | x-ai | 2,000,000 | $1.25 | $2.5 |
| SpaceXAI: Grok 4.20 | x-ai | 2,000,000 | $1.25 | $2.5 |
| Auto Router | openrouter | 2,000,000 | $-1,000,000 | $-1,000,000 |
| DeepSeek V4 Flash Latest | ~deepseek | 1,310,720 | $0.04 | $0.08 |
| DeepSeek: DeepSeek V4 Flash 0731 | deepseek | 1,310,720 | $0.14 | $0.28 |
| Meta: Llama 4 Scout | meta-llama | 1,310,720 | $0.1 | $0.3 |
| OpenAI: GPT-5.6 Luna Pro | openai | 1,050,000 | $0.2 | $1.2 |
| OpenAI: GPT-5.6 Luna Pro (batch) | openai | 1,050,000 | $0.1 | $0.6 |
| OpenAI: GPT-5.6 Luna | openai | 1,050,000 | $0.2 | $1.2 |
| OpenAI: GPT-5.6 Luna (batch) | openai | 1,050,000 | $0.1 | $0.6 |
| OpenAI: GPT-5.6 Terra Pro | openai | 1,050,000 | $2 | $12 |
| OpenAI: GPT-5.6 Terra Pro (batch) | openai | 1,050,000 | $1 | $6 |
| OpenAI: GPT-5.6 Terra | openai | 1,050,000 | $2 | $12 |
| OpenAI: GPT-5.6 Terra (batch) | openai | 1,050,000 | $1 | $6 |
| OpenAI: GPT-5.6 Sol Pro | openai | 1,050,000 | $2 | $10 |
| OpenAI: GPT-5.6 Sol Pro (batch) | openai | 1,050,000 | $1 | $5 |
| OpenAI: GPT-5.6 Sol | openai | 1,050,000 | $2 | $10 |
| OpenAI: GPT-5.6 Sol (batch) | openai | 1,050,000 | $1 | $5 |
最近变化
变化追踪将从下一次每日快照开始,排名与分数的变动会列在这里。
数据来源与方法
排行榜汇总自开放的社区评测。Elo 分数、置信区间与投票数来自 LMArena(社区投票的模型对战,以开放数据集发布)。价格与上下文元数据来自 OpenRouter 公开 API。Wikiprompt 每日快照并保留历史,不进行自有评测。