Why Blind Comparison Beats Rubrics: Designing the Harsh Critic
Three deliberate choices keep a gauntlet honest: a critic with fresh context, a binary pick instead of a score, and a bar that is a real thing instead of words.

Why Blind Comparison Beats Rubrics: Designing the Harsh Critic
Every failed self-improvement loop fails the same way: the model grades its own work, the grade inflates, the loop declares victory early. The gauntlet loop's answer is a critic designed around three deliberate choices. If you get these right, everything else is detail.
Choice 1: a different agent, fresh context
The critic must not be the builder wearing a different hat. It runs with fresh context: it has never seen the builder's plan, its excuses, or how many rounds it took. It sees two artifacts. This is not a courtesy; models are systematically kinder to work they produced, and kinder still to work they know was hard. Fresh context deletes that bias at the root.
Choice 2: a binary pick, not a score
Ask for a score out of 10 and you get 7 the first round, 8 the next, then 9, regardless of real progress. Scores are relative to expectations, and expectations move. A blind side-by-side pick cannot drift: the critic looks at yours and the bar, labels stripped, and says which is better. One bit of output, impossible to flatter.
Choice 3: the bar is a real thing, not words
A rubric is text the agent gets to interpret; interpretation always bends toward passing. A bar is a named, fetchable artifact: the actual Call of Duty, the actual competitor page, the actual published essay. The critic does not interpret quality, it observes defeat. When the work finally wins the blind pick, that means something.
The harshness is functional
"Make the critic really harsh" reads like flavor, but it is load-bearing: a neutral critic converges on approval because approval ends the task. Harshness plus one required output (name the single biggest gap) turns the critic into a gradient: every round, the builder knows exactly what to fix and nothing else.
Lineage
This is the evaluator-optimizer pattern from Anthropic's building-effective-agents playbook, pushed to its honest extreme - see From the Evaluator Pattern to the Gauntlet Loop. To see it in the wild: the original prompt, and the failure modes when one of the three choices is skipped.
Related Articles
- The Best Wan 2.1 Prompts: Open-Source AI Video That Works
Sep 3, 2026 · 7 min read
- Les Meilleurs Prompts Wan 2.1 : Vidéo IA Open-Source Qui Fonctionne
Sep 3, 2026 · 7 min read
- 最佳的Wan 2.1提示词:开源的AI视频,真的能用
Sep 3, 2026 · 7 min read
- As Melhores Prompts do Wan 2.1: Vídeo de IA Open-Source Que Funciona
Sep 3, 2026 · 7 min read
- Las Mejores Instrucciones para Wan 2.1: Video IA de Código Abierto Que Funciona
Sep 3, 2026 · 7 min read