Agentic arena
From Wikiprompt, the free prompt encyclopedia
Community Elo leaderboard, aggregated daily from LMArena head-to-head battles. Overall arena first, then per-category breakdowns.
Last updated: 2026-08-24 · 50 models
| Rank | Model | Organization | Elo score |
|---|---|---|---|
| 1 | Claude Opus 5 (High) | anthropic | |
| 2 | Claude Opus 5 (Max) | anthropic | |
| 3 | Claude Fable 5 (High) | anthropic | |
| 4 | Kimi K3 (Max) | moonshot | |
| 5 | GPT 5.6 Sol (xHigh) | openai | |
| 6 | Claude Opus 4.8 (High) | anthropic | |
| 7 | GPT 5.5 (xHigh) | openai | |
| 8 | Claude Opus 4.7 (High) | anthropic | |
| 9 | GPT 5.5 (High) | openai | |
| 10 | Claude Opus 4.7 | anthropic | |
| 11 | Claude Sonnet 5 (High) | anthropic | |
| 12 | Claude Opus 4.6 | anthropic | |
| 13 | GPT 5.5 | openai | |
| 14 | DeepSeek V4 Pro (High) (0813) | deepseek | |
| 15 | Qwen3.8 Max | alibaba | |
| 16 | Grok 4.5 | xai | |
| 17 | GLM 5.2 (Max) | zai | |
| 18 | GPT 5.4 (High) | openai | |
| 19 | GPT 5.6 Luna (xHigh) | openai | |
| 20 | Deepseek V4 Flash (High) (20260731) | deepseek | |
| 21 | Gemini 3.7 Flash (High) | ||
| 22 | GPT 5.6 Terra (xHigh) | openai | |
| 23 | Claude Sonnet 4.6 | anthropic | |
| 24 | Claude Opus 4.8 | anthropic | |
| 25 | Muse Spark 1.1 | meta | |
| 26 | Kimi K2.7 Code | moonshot | |
| 27 | DeepSeek V4 Pro | deepseek | |
| 28 | GLM 5.1 | zai | |
| 29 | Gemini 3.5 Flash (High) | ||
| 30 | Qwen3.7 Max | alibaba | |
| 31 | Gemini 3.1 Pro Preview | ||
| 32 | Kimi K2.6 | moonshot | |
| 33 | Mimo V2.5 Pro | xiaomi | |
| 34 | Hy3 | tencent | |
| 35 | Qwen3.7 Plus | alibaba | |
| 36 | Gemini 3.6 Flash (High) | ||
| 37 | Minimax M3 | minimax | |
| 38 | Gemini 3.5 Flash (Medium) | ||
| 39 | Inkling Small | thinky | |
| 40 | Inkling | thinky | |
| 41 | Mistral Medium 3.5 | mistral | |
| 42 | Grok 4.3 (High) | xai | |
| 43 | Gemini 3 Flash | ||
| 44 | Grok Build 0.1 | xai | |
| 45 | Solar Pro 4 | upstage | |
| 46 | Gemini 3.5 Flash Lite | ||
| 47 | Minimax M2.7 | minimax | |
| 48 | Nemotron 3 Ultra | nvidia | |
| 49 | Grok 4.3 | xai | |
| 50 | Gemma 4 31B |
See also
Sources and methodology
Rankings above are aggregated from open community evaluations. Elo ratings, confidence intervals and vote counts come from LMArena (crowdsourced head-to-head model battles, published as an open dataset). Model pricing and context metadata come from the public OpenRouter API. Wikiprompt snapshots these sources daily and preserves the history; we run no proprietary evaluations of our own.