LLM Benchmarks

From Wikiprompt, the free prompt encyclopedia

Live leaderboards for large language models and generative AI, aggregated daily from open, community-driven evaluations. Every table cites its source, and because Wikiprompt stores a snapshot every day, changes in the rankings stay on the record.

Last updated: 2026-08-24 · This is the LIVE leaderboard; dated snapshots preserve every past ranking. History →

Interactive explorer

Filter by arena, organization or license, search models, sort the table, and read the cost vs performance chart. Server tables below remain the citable reference.

Text arena

Click rows or chart points to pin models and compare only those.

Cost vs performance

Output price in USD per 1M tokens (log scale, from OpenRouter) vs community Elo (LMArena). Top-left = best value.

140014501500$0.01$0.1$1$10$100Output price, $ per 1M tokens (log)Community Elo

Elo over time (top text models)

claude-fable-5claude-opus-4-6claude-opus-4-6-highclaude-opus-4-7claude-opus-4-7-highclaude-opus-5-highgemini-3.1-pro-previewgemini-3.5-flash-high08-2308-24
Rank â–²ModelOrganizationElo scoreVotesInput $/1MOutput $/1MContext
1claude-opus-5-highanthropic1,50427,610$10$501,000,000
3claude-opus-4-6-highanthropic1,50372,359$5$251,000,000
4claude-opus-4-6anthropic1,49776,362$5$251,000,000
5claude-fable-5anthropic1,49524,331$10$501,000,000
7claude-opus-4-7-highanthropic1,49060,203$5$251,000,000
9claude-opus-4-7anthropic1,48361,304$5$251,000,000
10gemini-3.5-flash-highgoogle1,48330,004$1.5$91,048,576
12gemini-3.1-pro-previewgoogle1,48099,182$2$121,048,576
13gemini-3-progoogle1,47941,491$2$12131,072
14muse-spark-1.1meta1,47820,242$1.25$4.251,048,576
18gemini-3.5-flash-mediumgoogle1,47528,343$1.5$91,048,576
21qwen3.5-max-previewalibaba1,47121,484
22gpt-5.5-highopenai1,47159,545$5$301,050,000
23gpt-5.4-highopenai1,47060,708$2.5$151,050,000
24ernie-5.1baidu1,46837,137
25gemini-3-flashgoogle1,46630,852$0.5$31,048,576
26gpt-5.5openai1,46660,896$5$301,050,000
27glm-5.2-maxzai1,46530,637$0.97$3.041,048,576
28mimo-v2.5-proxiaomi1,46554,177$0.44$0.871,050,000
29glm-5.1zai1,46442,267$0.97$3.04204,800
30claude-opus-4-8-highanthropic1,46144,676$5$251,000,000
31claude-sonnet-4-6anthropic1,45866,512$2$101,000,000
32gemini-2.5-progoogle1,457124,807$1.25$101,048,576
33qwen3.7-plusalibaba1,45634,426$0.32$1.281,000,000
34kimi-k2.6moonshot1,45537,543$0.95$4262,144
35gpt-5.6-sol-xhighopenai1,45419,541$2$101,050,000
36gpt-5.4openai1,45363,659$2.5$151,050,000
37grok-4.5xai1,45222,030$2$6500,000
38claude-opus-4-8anthropic1,45245,331$5$251,000,000
39grok-4.20-beta-0309-reasoningxai1,45162,272$1.25$2.52,000,000
40deepseek-v4-prodeepseek1,45154,263$0.52$1.041,048,576
41grok-4.20-multi-agent-beta-0309xai1,45060,864$1.25$2.52,000,000
42claude-opus-4-5-20251101anthropic1,45070,985$5$251,000,000
43dola-seed-2.0-probytedance1,44874,513
45claude-opus-4-5-20251101-high-32kanthropic1,44737,198$5$251,000,000
47gpt-5.6-terra-xhighopenai1,44620,216$2$121,050,000
48glm-5zai1,44527,850$0.6$1.92204,800
49kimi-k2.5-thinkingmoonshot1,44570,943$0.45$2.25262,144
50deepseek-v4-pro-high-previewdeepseek1,44551,642$0.52$1.041,048,576
52ernie-5.0-0110baidu1,44435,267
53grok-4.20-beta1xai1,44426,770$1.25$2.52,000,000
56gemini-3-flash (thinking-minimal)google1,44286,359$0.5$31,048,576
57claude-sonnet-5-highanthropic1,44227,366$2$101,000,000
59gpt-5.1-highopenai1,44140,945$1.25$10400,000
62glm-4.6zai1,44035,830$0.5$2204,800
65qwen3-max-previewalibaba1,43927,838$0.78$3.9262,144
66gpt-5.2-chat-latest-20260210openai1,43934,339$1.75$14128,000
67qwen3.5-397b-a17balibaba1,43870,207$0.5$3.6262,144
68claude-sonnet-4-5-20250929anthropic1,43880,858$2$101,000,000
69grok-4.1-thinkingxai1,43765,550
70qwen3.6-plusalibaba1,43745,376$0.33$1.951,000,000
72mimo-v2-proxiaomi1,43724,397
74grok-4.1xai1,43667,767
76minimax-m3minimax1,43541,023$0.3$1.21,048,576
78claude-sonnet-4-5-20250929-high-32kanthropic1,43382,473$2$101,000,000
79deepseek-v4-flashdeepseek1,43249,001$0.04$0.081,310,720
80glm-4.5zai1,42924,403$0.6$2.2131,072
82chatgpt-4o-latest-20250326openai1,42982,766
83gpt-5.6-luna-xhighopenai1,42920,693$0.2$1.21,050,000
84mistral-large-3mistral1,42862,644$0.5$1.5262,144

Prices are matched to the closest OpenRouter listing by normalized model name; reasoning-effort variants share the base model price.

Takeaways
  • • Current text-arena leader: claude-opus-5-high (anthropic, Elo 1504)
  • • Best value (Elo per dollar, under $2/1M output): mimo-v2.5-pro (Elo 1465, $0.87/1M)
  • • Strongest open-license model: glm-5.2-max (MIT, #27)
  • • Coding leader (Aider pass rate): gpt-5 (high) (88%)

Text arena25 modelsFull leaderboard →

RankModelOrganizationElo score
1claude-opus-5-highanthropic1,504
3claude-opus-4-6-highanthropic1,503
4claude-opus-4-6anthropic1,497
5claude-fable-5anthropic1,495
7claude-opus-4-7-highanthropic1,490
9claude-opus-4-7anthropic1,483
10gemini-3.5-flash-highgoogle1,483
12gemini-3.1-pro-previewgoogle1,480
13gemini-3-progoogle1,479
14muse-spark-1.1meta1,478
18gemini-3.5-flash-mediumgoogle1,475
21qwen3.5-max-previewalibaba1,471
22gpt-5.5-highopenai1,471
23gpt-5.4-highopenai1,470
24ernie-5.1baidu1,468
25gemini-3-flashgoogle1,466
26gpt-5.5openai1,466
27glm-5.2-maxzai1,465
28mimo-v2.5-proxiaomi1,465
29glm-5.1zai1,464
30claude-opus-4-8-highanthropic1,461
31claude-sonnet-4-6anthropic1,458
32gemini-2.5-progoogle1,457
33qwen3.7-plusalibaba1,456
34kimi-k2.6moonshot1,455

Web development arena25 modelsFull leaderboard →

RankModelOrganizationElo score
1claude-opus-5-maxanthropic1,691
2kimi-k3-maxmoonshot1,674
3qwen3.8-maxalibaba1,669
4claude-opus-5-highanthropic1,663
5grok-4.6-highxai1,629
6claude-fable-5anthropic1,626
7gpt-5.6-sol-xhigh (codex-harness)openai1,619
8glm-5.3-maxzai1,599
9qwen3.8-27balibaba1,595
10gemini-3.7-flash-highgoogle1,587
11glm-5.2-maxzai1,582
12deepseek-v4-pro-high-20260813deepseek1,582
13deepseek-v4-flash-highdeepseek1,579
14claude-opus-4-8-highanthropic1,563
15claude-opus-4-7anthropic1,558
16claude-opus-4-7-highanthropic1,557
17grok-4.5xai1,556
18claude-opus-4-6-highanthropic1,546
19claude-opus-4-8anthropic1,539
20muse-spark-1.1meta1,539
21gemini-3.6-flash-highgoogle1,539
22claude-sonnet-5-highanthropic1,539
23claude-opus-4-6anthropic1,536
24muse-spark-1.2 (xHigh)meta1,534
25claude-sonnet-4-6anthropic1,522

Vision arena25 modelsFull leaderboard →

RankModelOrganizationElo score
1claude-fable-5anthropic1,328
2claude-opus-5-highanthropic1,323
3claude-opus-4-7anthropic1,317
4claude-opus-4-7-highanthropic1,316
5qwen3.8-maxalibaba1,315
6claude-opus-4-6-highanthropic1,315
7gemini-3.5-flash-highgoogle1,313
8claude-opus-4-6anthropic1,311
9gemini-3.5-flash-mediumgoogle1,308
10muse-sparkmeta1,306
11gemini-3-progoogle1,305
12muse-spark-1.2 (xHigh)meta1,304
13gemini-3.6-flash-highgoogle1,302
14gpt-5.4-highopenai1,297
15gpt-5.5openai1,295
16muse-spark-1.1meta1,295
17gemini-3.1-pro-previewgoogle1,294
18claude-opus-4-8-highanthropic1,293
19gpt-5.5-highopenai1,293
20gpt-5.4openai1,293
21grok-4.5xai1,291
22claude-opus-4-8anthropic1,287
23gemini-3-flashgoogle1,283
24claude-sonnet-4-6anthropic1,281
25claude-sonnet-5-highanthropic1,281

Text to image arena25 modelsFull leaderboard →

RankModelOrganizationElo score
1gpt-image-2 (medium)openai1,380
2reve-2.1reve1,302
3muse-imagemeta1,283
4reve-2.0reve1,270
5gemini-3.1-flash-image (nano-banana-2) [web-search]google1,263
6qwen-image-3.0-proalibaba1,258
7seedream-5.0-probytedance1,257
8mai-image-2.5microsoft-ai1,256
9gemini-3.1-flash-lite-image (nano-banana-2-lite)google1,251
10gemini-3-pro-image-2k (nano-banana-pro)google1,246
11gpt-image-1.5-high-fidelityopenai1,239
12gemini-3-pro-image-preview (nano-banana-pro)google1,232
13grok-imagine-image-qualityxai1,228
14ideogram-4.0-qualityideogram1,206
15qwen-image-2.0-pro-2026-06-22alibaba1,191
16uni-1.1-maxluma-ai1,188
17mai-image-2microsoft-ai1,182
18Cosmos3-Super-Text2Image (Agentic)nvidia1,181
19uni-1.1luma-ai1,180
20grok-imagine-imagexai1,172
21recraft-v4.1-utility-prorecraft1,169
22flux-2-maxbfl1,162
23grok-imagine-image-proxai1,161
24Cosmos3-Super-Text2Imagenvidia1,158
25flux-2-flexbfl1,156

Image editing arena25 modelsFull leaderboard →

RankModelOrganizationElo score
1gpt-image-2 (medium)openai1,463
2grok-imagine-image-2.0 (low)xai1,439
3mai-image-2.6-previewmicrosoft-ai1,420
4muse-imagemeta1,406
5mai-image-2.5microsoft-ai1,401
6seedream-5.0-probytedance1,394
7grok-imagine-image-quality (20260519)xai1,390
8gemini-3-pro-image-2k (nano-banana-pro)google1,390
9chatgpt-image-latest-high-fidelity (20251216)openai1,389
10gemini-3.1-flash-image (nano-banana-2) [web-search]google1,386
11gemini-3-pro-image-preview (nano-banana-pro)google1,385
12reve-2.1reve1,375
13gpt-image-1.5-high-fidelityopenai1,370
14reve-2.0reve1,358
15uni-1.1-maxluma-ai1,334
16grok-imagine-imagexai1,330
17gemini-3.1-flash-lite-image (nano-banana-2-lite)google1,314
18uni-1.1luma-ai1,312
19qwen-image-2.0-pro-2026-06-22alibaba1,304
20hunyuan-image-3.0-instructtencent1,303
21wan2.7-image-proalibaba1,302
22seedream-4.5bytedance1,302
23wan2.7-imagealibaba1,301
24gemini-2.5-flash-image-preview (nano-banana)google1,293
25seedream-5.0-litebytedance1,293

Text to video arena25 modelsFull leaderboard →

RankModelOrganizationElo score
1gemini-omni-flashgoogle1,512
2flux-3-videobfl1,494
3dreamina-seedance-2.0-720pbytedance1,482
4dreamina-seedance-2.5-720pbytedance1,477
5muse-videometa1,457
6minimax-h3minimax1,453
7happyhorse-1.0aorizon1,428
8sora-2-proopenai1,364
9veo-3.1-audiogoogle1,364
10veo-3.1-audio-1080pgoogle1,363
11veo-3.1-fast-audiogoogle1,361
12veo-3.1-fast-audio-1080pgoogle1,358
13veo-3-fast-audiogoogle1,347
14grok-imagine-video-720pxai1,345
15wan2.7-t2vwan1,344
16sora-2openai1,340
17veo-3-audiogoogle1,339
18wan2.6-t2valibaba1,330
19seedance-v1.5-probytedance1,256
20veo-3google1,253
21wan2.5-t2v-previewalibaba1,249
22veo-3-fastgoogle1,248
23pixverse-v5.61,240
24runway-gen-4.5runway1,223
25kling-2.5-turbo-1080pkling1,219

Video editing arena9 modelsFull leaderboard →

RankModelOrganizationElo score
1dreamina-seedance-2.5-720pbytedance1,411
2minimax-h3minimax1,388
3dreamina-seedance-2.0-720pbytedance1,359
4gemini-omni-flashgoogle1,358
5happyhorse-1.0aorizon1,307
6grok-imagine-videoxai1,262
7kling-o3-prokling1,257
8kling-o1-prokling1,200
9runway-gen4-alephrunway1,183

Agentic arena25 modelsFull leaderboard →

RankModelOrganizationElo score
1Claude Opus 5 (High)anthropic
2Claude Opus 5 (Max)anthropic
3Claude Fable 5 (High)anthropic
4Kimi K3 (Max)moonshot
5GPT 5.6 Sol (xHigh)openai
6Claude Opus 4.8 (High)anthropic
7GPT 5.5 (xHigh)openai
8Claude Opus 4.7 (High)anthropic
9GPT 5.5 (High)openai
10Claude Opus 4.7anthropic
11Claude Sonnet 5 (High)anthropic
12Claude Opus 4.6anthropic
13GPT 5.5openai
14DeepSeek V4 Pro (High) (0813)deepseek
15Qwen3.8 Maxalibaba
16Grok 4.5xai
17GLM 5.2 (Max)zai
18GPT 5.4 (High)openai
19GPT 5.6 Luna (xHigh)openai
20Deepseek V4 Flash (High) (20260731)deepseek
21Gemini 3.7 Flash (High)google
22GPT 5.6 Terra (xHigh)openai
23Claude Sonnet 4.6anthropic
24Claude Opus 4.8anthropic
25Muse Spark 1.1meta

Document arena25 modelsFull leaderboard →

RankModelOrganizationElo score
1claude-opus-5-highanthropic1,520
2claude-opus-4-6anthropic1,510
3claude-opus-4-6-thinkinganthropic1,506
4claude-fable-5anthropic1,504
5claude-opus-4-7anthropic1,498
6claude-opus-4-7-thinkinganthropic1,497
7gpt-5.5-highopenai1,485
8claude-sonnet-4-6anthropic1,483
9gpt-5.5openai1,480
10gpt-5.6-terra-xhighopenai1,479
11gpt-5.6-sol-xhighopenai1,479
12claude-opus-4-8-thinkinganthropic1,475
13muse-spark-1.1meta1,472
14claude-sonnet-5-highanthropic1,470
15gpt-5.4openai1,470
16claude-opus-4-8anthropic1,469
17gemini-3.5-flash-mediumgoogle1,465
18gpt-5.6-luna-xhighopenai1,462
19claude-opus-4-5-20251101anthropic1,462
20grok-4.5xai1,454
21kimi-k2.6moonshot1,451
22claude-sonnet-4-5-20250929anthropic1,446
23gemini-3.1-pro-previewgoogle1,445
24muse-sparkmeta1,443
25qwen3.7-plusalibaba1,440

Other leaderboards

Aider polyglot (coding)

Metric: pass rate 2 (%) · aider.chat/docs/leaderboards/

RankModelOrganizationScore
1gpt-5 (high)88
2gpt-5 (medium)86.7
3o3-pro (high)84.9
4gemini-2.5-pro-preview-06-05 (32k think)83.1
5o3 (high)81.3
6gpt-5 (low)81.3
7grok-4 (high)79.6
8gemini-2.5-pro-preview-06-05 (default think)79.1
9o3 (high) + gpt-4.178.2
10Gemini 2.5 Pro Preview 05-0676.9
11o376.9
12DeepSeek-V3.2-Exp (Reasoner)74.2
13Gemini 2.5 Pro Preview 03-2572.9
14o4-mini (high)72
15claude-opus-4-20250514 (32k thinking)72

LiveBench

Metric: average score (0-100) · livebench.ai

RankModelOrganizationScore
1claude-fable-5-max-effort83.4
2gpt-5.6-sol-max81.7
3gpt-5.5-xhigh80.8
4claude-opus-5-max-effort80.5
5gemini-3.7-flash-high79.9
6smaug-agentic79.7
7qwen3.8-max79.5
8kimi-k379.5
9grok-4.679
10muse-spark-1.2-xhigh78.9
11gpt-5.4-xhigh78.8
12gpt-5.6-terra-max78.6
13deepseek-v4-pro-081378.2
14gemini-3.1-pro-preview-high78
15deepseek-v4-flash-vision-exp77.7

BenchLM

Metric: aggregated display score · benchlm.ai

RankModelOrganizationScore
1Claude Mythos 5Anthropic83
2Claude Opus 5Anthropic82.7
3Claude Fable 5Anthropic82.7
4GPT-5.6 SolOpenAI81.7
5Kimi K3Moonshot AI80.2
6Qwen3.8 MaxAlibaba79
7Muse Spark 1.1Meta76.7
8Claude Opus 4.8Anthropic76.2
9Gemini 3.6 FlashGoogle75.2
10Grok 4.5xAI75.2
11GPT-5.5OpenAI73.3
12GPT-5.4OpenAI73.1
13GPT-5.6 TerraOpenAI72.6
14Claude Opus 4.7 (Adaptive)Anthropic72.3
15Claude Opus 4.7Anthropic71.9

Artificial Analysis (Intelligence Index)

Metric: AA Intelligence Index (0-100) · artificialanalysis.ai

RankModelOrganizationScore
1Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic63.1
2Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic62.5
3Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic62.1
4Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic61.5
5GPT-5.6 Sol (max)OpenAI60.9
6Grok 4.6 (high)SpaceXAI60.9
7Grok 4.6 (xhigh)SpaceXAI60
8Kimi K3 (max)Kimi59.7
9GLM-5.3 (max)Z AI59.5
10GPT-5.6 Sol (xhigh)OpenAI59
11Grok 4.6 (medium)SpaceXAI59
12Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic58.6
13Qwen3.8 MaxAlibaba58.1
14Qwen3.8 2.4T A95BAlibaba57.7
15GPT-5.6 Sol (high)OpenAI57.3

Pricing and context windows

Model pricing and context metadata from the public OpenRouter API. Prices are USD per 1M tokens.

ModelOrganizationContextInput $/1MOutput $/1M
Auto Router (Beta)openrouter2,000,000$-1,000,000$-1,000,000
Pareto Code Routeropenrouter2,000,000$-1,000,000$-1,000,000
SpaceXAI: Grok 4.20 Multi-Agentx-ai2,000,000$1.25$2.5
SpaceXAI: Grok 4.20x-ai2,000,000$1.25$2.5
Auto Routeropenrouter2,000,000$-1,000,000$-1,000,000
DeepSeek V4 Flash Latest~deepseek1,310,720$0.04$0.08
DeepSeek: DeepSeek V4 Flash 0731deepseek1,310,720$0.14$0.28
Meta: Llama 4 Scoutmeta-llama1,310,720$0.1$0.3
OpenAI: GPT-5.6 Luna Proopenai1,050,000$0.2$1.2
OpenAI: GPT-5.6 Luna Pro (batch)openai1,050,000$0.1$0.6
OpenAI: GPT-5.6 Lunaopenai1,050,000$0.2$1.2
OpenAI: GPT-5.6 Luna (batch)openai1,050,000$0.1$0.6
OpenAI: GPT-5.6 Terra Proopenai1,050,000$2$12
OpenAI: GPT-5.6 Terra Pro (batch)openai1,050,000$1$6
OpenAI: GPT-5.6 Terraopenai1,050,000$2$12
OpenAI: GPT-5.6 Terra (batch)openai1,050,000$1$6
OpenAI: GPT-5.6 Sol Proopenai1,050,000$2$10
OpenAI: GPT-5.6 Sol Pro (batch)openai1,050,000$1$5
OpenAI: GPT-5.6 Solopenai1,050,000$2$10
OpenAI: GPT-5.6 Sol (batch)openai1,050,000$1$5

Recent changes

Change tracking starts with the next daily snapshot. Rank and score movements will be listed here.

Sources and methodology

Rankings above are aggregated from open community evaluations. Elo ratings, confidence intervals and vote counts come from LMArena (crowdsourced head-to-head model battles, published as an open dataset). Model pricing and context metadata come from the public OpenRouter API. Wikiprompt snapshots these sources daily and preserves the history; we run no proprietary evaluations of our own.

  1. LMArena leaderboard dataset (Hugging Face)
  2. LMArena leaderboard
  3. OpenRouter models API
  4. Artificial Analysis (intelligence, speed and latency data)

See also