BlogGuides

Introducing Wikiprompt LLM Benchmarks: Live Leaderboards, With Receipts

Wikiprompt now tracks LLM leaderboards across ten arenas, snapshotted daily so ranking changes stay on the record. Sourced from LMArena and OpenRouter, cited like Wikipedia references, in seven languages.

Introducing Wikiprompt LLM Benchmarks: Live Leaderboards, With Receipts

Introducing Wikiprompt LLM Benchmarks: Live Leaderboards, With Receipts

Wikiprompt now has a benchmarks section: wikiprompt.org/benchmarks. Live leaderboards for large language models and generative AI, covering ten arenas: text, web development, vision, text to image, image editing, text to video, video editing, search, agents and documents. Free, in seven languages, and updated every day.

Why another leaderboard site?

Because benchmark numbers move quietly. A model climbs three places, a score gets recalculated, yesterday's leader slides to fourth, and there is usually no record that anything changed. We treat benchmarks the way Wikipedia treats articles: every day Wikiprompt stores a full snapshot of the rankings, and the page shows recent changes against the previous snapshot. The history stays on the record.

The second reason is sourcing. Our tables cite where every number comes from, like Wikipedia references:

  • Elo ratings, confidence intervals and vote counts come from LMArena, the crowdsourced head-to-head model battles, published as an open dataset.
  • Pricing, context windows and knowledge cutoffs come from the public OpenRouter API.
  • We run no proprietary evaluations and we do not editorialize the scores. We aggregate open, community-driven data, cite it, and keep the receipts.

    What the arenas say right now

    A few current leaders, straight from today's snapshot:

  • Text: claude-opus-4-6-high leads with an Elo around 1527 across more than 45,000 community votes.
  • Text to image: gpt-image-2 (medium) tops the arena with roughly 69,000 votes; it also leads image editing with almost 200,000.
  • Text to video: gemini-omni-flash holds first place.
  • Video editing: dreamina-seedance-2.5 is the model to beat.
  • Search: claude-opus-4-6-search leads over 112,000 votes.
  • These will change. That is the point: when they do, the movement shows up in the Recent changes section instead of vanishing.

    How it connects to the prompt encyclopedia

    Wikiprompt already catalogs tens of thousands of prompts tagged by model. The benchmarks section closes the loop: see which model leads an arena, then jump to real prompts for that model to put it to work. Model guides like The Best GPT Image Prompts and the Nano Banana family comparison sit alongside the rankings.

    Open, like everything else here

    The benchmarks page follows the same principles as the rest of Wikiprompt: the compilation is CC BY-SA, the sources are open datasets and public APIs, and the whole prompt catalog is downloadable as a single file. If you want the benchmark data itself, the sources we cite are one click away.

    Check the current standings at wikiprompt.org/benchmarks. They will look different next week, and you will be able to see exactly how.

    Tags
    benchmarks·leaderboard·lmarena·llm·rankings·announcement