All benchmarks

Agentic arena

From Wikiprompt, the free prompt encyclopedia

Community Elo leaderboard, aggregated daily from LMArena head-to-head battles. Overall arena first, then per-category breakdowns.

Last updated: 2026-08-24 · 50 models

RankModelOrganizationElo score
1Claude Opus 5 (High)anthropic
2Claude Opus 5 (Max)anthropic
3Claude Fable 5 (High)anthropic
4Kimi K3 (Max)moonshot
5GPT 5.6 Sol (xHigh)openai
6Claude Opus 4.8 (High)anthropic
7GPT 5.5 (xHigh)openai
8Claude Opus 4.7 (High)anthropic
9GPT 5.5 (High)openai
10Claude Opus 4.7anthropic
11Claude Sonnet 5 (High)anthropic
12Claude Opus 4.6anthropic
13GPT 5.5openai
14DeepSeek V4 Pro (High) (0813)deepseek
15Qwen3.8 Maxalibaba
16Grok 4.5xai
17GLM 5.2 (Max)zai
18GPT 5.4 (High)openai
19GPT 5.6 Luna (xHigh)openai
20Deepseek V4 Flash (High) (20260731)deepseek
21Gemini 3.7 Flash (High)google
22GPT 5.6 Terra (xHigh)openai
23Claude Sonnet 4.6anthropic
24Claude Opus 4.8anthropic
25Muse Spark 1.1meta
26Kimi K2.7 Codemoonshot
27DeepSeek V4 Prodeepseek
28GLM 5.1zai
29Gemini 3.5 Flash (High)google
30Qwen3.7 Maxalibaba
31Gemini 3.1 Pro Previewgoogle
32Kimi K2.6moonshot
33Mimo V2.5 Proxiaomi
34Hy3tencent
35Qwen3.7 Plusalibaba
36Gemini 3.6 Flash (High)google
37Minimax M3minimax
38Gemini 3.5 Flash (Medium)google
39Inkling Smallthinky
40Inklingthinky
41Mistral Medium 3.5mistral
42Grok 4.3 (High)xai
43Gemini 3 Flashgoogle
44Grok Build 0.1xai
45Solar Pro 4upstage
46Gemini 3.5 Flash Litegoogle
47Minimax M2.7minimax
48Nemotron 3 Ultranvidia
49Grok 4.3xai
50Gemma 4 31Bgoogle

See also

Sources and methodology

Rankings above are aggregated from open community evaluations. Elo ratings, confidence intervals and vote counts come from LMArena (crowdsourced head-to-head model battles, published as an open dataset). Model pricing and context metadata come from the public OpenRouter API. Wikiprompt snapshots these sources daily and preserves the history; we run no proprietary evaluations of our own.

  1. LMArena leaderboard dataset (Hugging Face)
  2. LMArena leaderboard