Arene agents
De Wikiprompt, l’encyclopedie libre des prompts
Classement Elo communautaire, agrege chaque jour depuis les duels LMArena. L’arene generale d’abord, puis le detail par categorie.
Derniere mise a jour: 2026-08-24 · 50 modeles
| Rang | Modele | Organisation | Score Elo |
|---|---|---|---|
| 1 | Claude Opus 5 (High) | anthropic | |
| 2 | Claude Opus 5 (Max) | anthropic | |
| 3 | Claude Fable 5 (High) | anthropic | |
| 4 | Kimi K3 (Max) | moonshot | |
| 5 | GPT 5.6 Sol (xHigh) | openai | |
| 6 | Claude Opus 4.8 (High) | anthropic | |
| 7 | GPT 5.5 (xHigh) | openai | |
| 8 | Claude Opus 4.7 (High) | anthropic | |
| 9 | GPT 5.5 (High) | openai | |
| 10 | Claude Opus 4.7 | anthropic | |
| 11 | Claude Sonnet 5 (High) | anthropic | |
| 12 | Claude Opus 4.6 | anthropic | |
| 13 | GPT 5.5 | openai | |
| 14 | DeepSeek V4 Pro (High) (0813) | deepseek | |
| 15 | Qwen3.8 Max | alibaba | |
| 16 | Grok 4.5 | xai | |
| 17 | GLM 5.2 (Max) | zai | |
| 18 | GPT 5.4 (High) | openai | |
| 19 | GPT 5.6 Luna (xHigh) | openai | |
| 20 | Deepseek V4 Flash (High) (20260731) | deepseek | |
| 21 | Gemini 3.7 Flash (High) | ||
| 22 | GPT 5.6 Terra (xHigh) | openai | |
| 23 | Claude Sonnet 4.6 | anthropic | |
| 24 | Claude Opus 4.8 | anthropic | |
| 25 | Muse Spark 1.1 | meta | |
| 26 | Kimi K2.7 Code | moonshot | |
| 27 | DeepSeek V4 Pro | deepseek | |
| 28 | GLM 5.1 | zai | |
| 29 | Gemini 3.5 Flash (High) | ||
| 30 | Qwen3.7 Max | alibaba | |
| 31 | Gemini 3.1 Pro Preview | ||
| 32 | Kimi K2.6 | moonshot | |
| 33 | Mimo V2.5 Pro | xiaomi | |
| 34 | Hy3 | tencent | |
| 35 | Qwen3.7 Plus | alibaba | |
| 36 | Gemini 3.6 Flash (High) | ||
| 37 | Minimax M3 | minimax | |
| 38 | Gemini 3.5 Flash (Medium) | ||
| 39 | Inkling Small | thinky | |
| 40 | Inkling | thinky | |
| 41 | Mistral Medium 3.5 | mistral | |
| 42 | Grok 4.3 (High) | xai | |
| 43 | Gemini 3 Flash | ||
| 44 | Grok Build 0.1 | xai | |
| 45 | Solar Pro 4 | upstage | |
| 46 | Gemini 3.5 Flash Lite | ||
| 47 | Minimax M2.7 | minimax | |
| 48 | Nemotron 3 Ultra | nvidia | |
| 49 | Grok 4.3 | xai | |
| 50 | Gemma 4 31B |
Voir aussi
Sources et methodologie
Les classements sont agreges depuis des evaluations communautaires ouvertes. Les scores Elo, intervalles de confiance et votes proviennent de LMArena (duels de modeles votes par la communaute, publies en dataset ouvert). Les prix et metadonnees de contexte proviennent de l’API publique OpenRouter. Wikiprompt prend un instantane quotidien et preserve l’historique ; nous ne menons aucune evaluation proprietaire.