Blog›Benchmarks

LLM Benchmark Snapshot, September 22, 2026: A Frozen Week at the Top

Not one top-3 position changed across LiveBench, BenchLM, Aider and LMArena this week. The real action was off the boards: the Jev and Laya decision models opened a category no benchmark measures yet.

LLM Benchmark Snapshot, September 22, 2026: A Frozen Week at the Top

LLM Benchmark Snapshot, September 22, 2026: A Frozen Week at the Top

Every week we photograph the major LLM leaderboards and keep the receipts at /benchmarks/history. This week's photo, September 22, is notable for what did not happen: across all five sources we track, not a single top-3 position changed hands.

The leaders, by source

  • LiveBench (overall): Claude Fable 5.1 (max effort) holds first at 83.78, ahead of Fable 5 (83.38) and GPT-6 Astra max (83.05). Three models within 0.73 points.
  • BenchLM (overall): Claude Fable 5.1 at 84.58, GPT-6 Astra at 83.79, Claude Opus 5 at 81.87.
  • Aider polyglot (coding): GPT-5 (high) keeps its 88.0, with GPT-5 (medium) at 86.7 and o3-pro at 84.9. OpenAI still owns this board.
  • LMArena text Elo: Claude Fable 5.1 max at 1507.6, with two Opus 5 variants breathing down its neck at 1505 - a 2.6-Elo gap, effectively a tie.
  • LMArena webdev Elo: GPT-6 Astra max at 1800.3, comfortably ahead of Fable 5.1 max (1758) - the one board where the gap is wide.
  • The sources still disagree

    The disagreement pattern we flag every week holds: Anthropic leads general-purpose text on three boards, OpenAI leads coding (Aider) and webdev (LMArena) decisively. If you pick a model by "the leaderboard," the answer still depends entirely on which leaderboard matches your workload.

    The action was off the leaderboards

    A frozen week at the top does not mean a quiet week. The most-discussed launches were not chat models at all: TypeSafe's Jev and the open-source, Apache 2.0 Laya from India's Convai Innovations opened up the "System One" decision-model category - typed answers with calibrated probabilities instead of generated text. None of the boards above measure that kind of model yet, which is itself a gap worth watching: the benchmark ecosystem measures writing and coding, while a growing slice of production AI is fast classification. Our comparison: Laya vs Jev.

    Browse the full snapshot at /benchmarks/snapshot/2026-09-22, or the live board at /benchmarks.

    Tags
    benchmarks·llm·leaderboard·claude·gpt-6·aider·livebench·weekly-snapshot