Alle Benchmarks

Dokumenten-Arena

Von Wikiprompt, der freien Prompt-Enzyklopadie

Community-Elo-Rangliste, taglich aus LMArena-Duellen aggregiert. Zuerst die Gesamt-Arena, dann die Aufschlusselung nach Kategorie.

Zuletzt aktualisiert: 2026-08-24 · 38 Modelle

RangModellOrganisationElo-Wert
1claude-opus-5-highanthropic1,520
2claude-opus-4-6anthropic1,510
3claude-opus-4-6-thinkinganthropic1,506
4claude-fable-5anthropic1,504
5claude-opus-4-7anthropic1,498
6claude-opus-4-7-thinkinganthropic1,497
7gpt-5.5-highopenai1,485
8claude-sonnet-4-6anthropic1,483
9gpt-5.5openai1,480
10gpt-5.6-terra-xhighopenai1,479
11gpt-5.6-sol-xhighopenai1,479
12claude-opus-4-8-thinkinganthropic1,475
13muse-spark-1.1meta1,472
14claude-sonnet-5-highanthropic1,470
15gpt-5.4openai1,470
16claude-opus-4-8anthropic1,469
17gemini-3.5-flash-mediumgoogle1,465
18gpt-5.6-luna-xhighopenai1,462
19claude-opus-4-5-20251101anthropic1,462
20grok-4.5xai1,454
21kimi-k2.6moonshot1,451
22claude-sonnet-4-5-20250929anthropic1,446
23gemini-3.1-pro-previewgoogle1,445
24muse-sparkmeta1,443
25qwen3.7-plusalibaba1,440
26minimax-m3minimax1,435
27gemini-3-progoogle1,433
28kimi-k2.5-thinkingmoonshot1,429
29gemma-4-31bgoogle1,424
30claude-haiku-4-5-20251001anthropic1,423
31gemini-2.5-progoogle1,421
32glm-5v-turbozai1,417
33grok-4.20-beta-0309-reasoningxai1,414
34gemini-3-flashgoogle1,413
35gpt-5.2-highopenai1,405
36gpt-5.5-instantopenai1,402
37gpt-5.1openai1,401
38gpt-5.2openai1,400

Siehe auch

Quellen und Methodik

Die Ranglisten werden aus offenen Community-Bewertungen aggregiert. Elo-Werte, Konfidenzintervalle und Stimmen stammen von LMArena (Community-Votings in direkten Modell-Duellen, als offenes Dataset veroffentlicht). Preise und Kontext-Metadaten stammen aus der offentlichen OpenRouter-API. Wikiprompt erstellt taglich einen Snapshot und bewahrt die Historie; eigene Evaluationen fuhren wir nicht durch.

  1. LMArena leaderboard dataset (Hugging Face)
  2. LMArena leaderboard