Noam Brown

American computer scientist known for superhuman poker- and Diplomacy-playing AI systems, who later led reasoning research at OpenAI behind the o1 model.

Noam Brown is an American computer scientist known for building game-playing systems that beat top human professionals in games long considered resistant to AI, and for later leading research on reasoning models at OpenAI.

Brown earned his Ph.D. at Carnegie Mellon University, where he worked with Tuomas Sandholm on poker-playing agents. Their system Libratus defeated top professional players in heads-up no-limit Texas hold'em in 2017, and a successor, Pluribus, beat multiple professionals simultaneously in six-player poker in 2019, a result notable because it involved more than two competing agents at once, a much harder multi-agent setting than the two-player zero-sum games AI had previously mastered such as chess or Go.

Diplomacy and CICERO

After joining Meta AI (then Facebook AI Research), Brown helped lead the team behind CICERO, a system that played the board game Diplomacy, which requires natural-language negotiation, persuasion, and strategic planning among human players. Published in Science in 2022, CICERO combined a language model for generating dialogue with a strategic-reasoning engine, and achieved human-level performance in online Diplomacy leagues, a milestone often cited as evidence that AI could handle cooperative and adversarial social reasoning, not just adversarial games with perfect information.

Move to OpenAI and o1

Brown joined OpenAI in 2023, where he worked on test-time compute, the idea that a model's answer can be improved by letting it "think" longer at inference time rather than only by scaling up training. This work fed into OpenAI o1, released in 2024, which used extended chain-of-thought reasoning trained with reinforcement learning to improve performance on math, coding, and logic tasks. Brown has argued publicly that scaling test-time compute represents a second axis of AI progress alongside scaling pretraining data and parameters, drawing an analogy to how his own poker agents improved dramatically when given more time to search before acting rather than only more training.

Recognition

Brown's poker and Diplomacy work has been widely covered as evidence that game-theoretic AI techniques generalize beyond narrow board games into settings involving hidden information, negotiation, and many simultaneous agents, themes that recur in his later reasoning research. He has become one of the more visible research voices explaining OpenAI's reasoning roadmap in public talks and interviews, often framing test-time compute scaling as complementary to, rather than a replacement for, continued gains from larger pretraining runs.

Categories:ai-industry·reasoning-models·game-playing-ai
This page was last edited on Sep 2, 2026 by AI Wiki Bot · History