Arene documents
De Wikiprompt, l’encyclopedie libre des prompts
Classement Elo communautaire, agrege chaque jour depuis les duels LMArena. L’arene generale d’abord, puis le detail par categorie.
Derniere mise a jour: 2026-08-24 · 38 modeles
| Rang | Modele | Organisation | Score Elo |
|---|---|---|---|
| 1 | claude-opus-5-high | anthropic | 1,520 |
| 2 | claude-opus-4-6 | anthropic | 1,510 |
| 3 | claude-opus-4-6-thinking | anthropic | 1,506 |
| 4 | claude-fable-5 | anthropic | 1,504 |
| 5 | claude-opus-4-7 | anthropic | 1,498 |
| 6 | claude-opus-4-7-thinking | anthropic | 1,497 |
| 7 | gpt-5.5-high | openai | 1,485 |
| 8 | claude-sonnet-4-6 | anthropic | 1,483 |
| 9 | gpt-5.5 | openai | 1,480 |
| 10 | gpt-5.6-terra-xhigh | openai | 1,479 |
| 11 | gpt-5.6-sol-xhigh | openai | 1,479 |
| 12 | claude-opus-4-8-thinking | anthropic | 1,475 |
| 13 | muse-spark-1.1 | meta | 1,472 |
| 14 | claude-sonnet-5-high | anthropic | 1,470 |
| 15 | gpt-5.4 | openai | 1,470 |
| 16 | claude-opus-4-8 | anthropic | 1,469 |
| 17 | gemini-3.5-flash-medium | 1,465 | |
| 18 | gpt-5.6-luna-xhigh | openai | 1,462 |
| 19 | claude-opus-4-5-20251101 | anthropic | 1,462 |
| 20 | grok-4.5 | xai | 1,454 |
| 21 | kimi-k2.6 | moonshot | 1,451 |
| 22 | claude-sonnet-4-5-20250929 | anthropic | 1,446 |
| 23 | gemini-3.1-pro-preview | 1,445 | |
| 24 | muse-spark | meta | 1,443 |
| 25 | qwen3.7-plus | alibaba | 1,440 |
| 26 | minimax-m3 | minimax | 1,435 |
| 27 | gemini-3-pro | 1,433 | |
| 28 | kimi-k2.5-thinking | moonshot | 1,429 |
| 29 | gemma-4-31b | 1,424 | |
| 30 | claude-haiku-4-5-20251001 | anthropic | 1,423 |
| 31 | gemini-2.5-pro | 1,421 | |
| 32 | glm-5v-turbo | zai | 1,417 |
| 33 | grok-4.20-beta-0309-reasoning | xai | 1,414 |
| 34 | gemini-3-flash | 1,413 | |
| 35 | gpt-5.2-high | openai | 1,405 |
| 36 | gpt-5.5-instant | openai | 1,402 |
| 37 | gpt-5.1 | openai | 1,401 |
| 38 | gpt-5.2 | openai | 1,400 |
Voir aussi
Sources et methodologie
Les classements sont agreges depuis des evaluations communautaires ouvertes. Les scores Elo, intervalles de confiance et votes proviennent de LMArena (duels de modeles votes par la communaute, publies en dataset ouvert). Les prix et metadonnees de contexte proviennent de l’API publique OpenRouter. Wikiprompt prend un instantane quotidien et preserve l’historique ; nous ne menons aucune evaluation proprietaire.