BlogBenchmarks

LLM Benchmark Snapshot: August 30, 2026

Weekly frozen photo of six LLM leaderboards: Opus 5 max takes the text arena, Fable 5 holds LiveBench, gpt-5 keeps Aider coding, and the sources still disagree.

LLM Benchmark Snapshot: August 30, 2026

LLM Benchmark Snapshot: August 30, 2026

Every Sunday Wikiprompt freezes a photo of the major LLM leaderboards. This is the snapshot for the week ending August 30, 2026 - 14,164 entries across six sources, browsable in full at /benchmarks/snapshot/2026-08-30.

Leaders by source

  • LMArena (text): claude-opus-5-max takes the top Elo (1504.7), a hair over claude-opus-5-high (1503.6), with the previous-generation claude-opus-4-6-high still holding third.
  • LMArena (web development): claude-opus-5-max leads by a wide margin (1690.6), with kimi-k3-max (1673.7) and qwen3.8-max (1668.9) as the strongest open-weight challengers.
  • LiveBench: claude-fable-5-max-effort on top (83.4), ahead of gpt-5.6-sol-max (81.7) and gpt-5.5-xhigh (80.8).
  • BenchLM: the Claude 5 family sweeps the podium - Mythos 5 (83.4), Fable 5 (83.2), Opus 5 (83.1) - separated by fractions of a point.
  • Aider polyglot (coding): gpt-5 (high) remains the coding-benchmark king at 88.0% pass rate, ahead of gpt-5 (medium) and o3-pro (high).
  • Artificial Analysis intelligence index: Claude Opus 5 at max effort leads (63.1), with the Fable 5 variants right behind.
  • What changed vs last week

    The text-arena crown changed hands inside the same family: claude-opus-5-high (1504.2 last week) was edged out by claude-opus-5-max (1504.7). Statistically a coin flip, editorially a reminder that effort settings now matter as much as model choice at the top of the table.

    The sources still disagree

    As every week, no two leaderboards crown the same model: LMArena's community Elo favors Opus 5, LiveBench's contamination-free academic sets favor Fable 5, and Aider's real-world coding harness still belongs to gpt-5. Cross-reference before you choose - that is what the comparison view is for.

    Browse the full dated table at /benchmarks/snapshot/2026-08-30, or the whole archive at /benchmarks/history.

    Tags
    benchmarks·llm·leaderboard·weekly