Humanity's Last Exam is a benchmark dataset designed to evaluate the capabilities of advanced artificial intelligence systems. It consists of a curated set of highly challenging questions, authored by experts across numerous disciplines, intended to probe the limits of current large language models and related technologies. The benchmark was created to provide a rigorous, forward-looking measure of AI progress, particularly as models approach or surpass human performance on existing evaluation suites.
The project was conceived as a collaborative effort involving researchers and practitioners from leading AI organizations, including OpenAI and other prominent institutions. It was publicly released in early 2025, with the goal of establishing a new standard for assessing expert-level knowledge and reasoning in AI systems. The benchmark's name reflects its ambition to represent a final, comprehensive test of AI's intellectual abilities, at least for the current generation of models.
Design and Structure
The benchmark comprises thousands of multiple-choice and short-answer questions, each crafted by subject-matter experts. The questions span a wide range of fields, including mathematics, physics, chemistry, biology, computer science, history, law, and the humanities. A key design principle is that the questions should be difficult for even highly capable AI systems, but also verifiable and unambiguous in their correct answers. To ensure quality, each submitted question undergoes a rigorous review process, and questions that can be solved by existing models are typically rejected.
The dataset is divided into a public test set and a private, held-out set. The private set is used for official evaluations to prevent overfitting and data contamination. The public set allows researchers and developers to benchmark their own systems, while the private set provides a more reliable measure of generalizable capability. The benchmark's creators also implemented strict protocols to prevent models from being trained on the test data, including monitoring for memorization and requiring that all questions be novel and not publicly available prior to release.
Key Participants and Contributors
The initiative was led by a team of researchers, including Jacob Steinhardt and David Kaplan, who are known for their work on AI safety and evaluation. They were joined by contributors from Anthropic, Google DeepMind, and other major AI labs, as well as academics from institutions such as MIT CSAIL, Stanford AI Lab, and BAIR (Berkeley AI Research). The effort also drew on the expertise of independent researchers and domain specialists from around the world, who submitted questions in their respective fields.
The collaborative nature of the project was a deliberate choice, aimed at ensuring diversity in question types and subject coverage. The organizers solicited contributions through public calls for questions, attracting submissions from thousands of experts. Each submission was reviewed by multiple independent evaluators to verify its accuracy, difficulty, and suitability. This process resulted in a final dataset that is both broad in scope and high in quality.
Performance of Current AI Systems
Initial results from the benchmark revealed that even the most advanced AI systems, including those based on Large language model architectures, scored well below human expert levels. The best-performing models, such as those developed by OpenAI and Google DeepMind, achieved accuracy rates in the single digits to low teens on the test set. This was in stark contrast to their performance on earlier benchmarks, where they often exceeded 90% accuracy.
The low scores were attributed to the benchmark's design, which specifically targets reasoning, multi-step problem-solving, and deep domain knowledge. Many questions require combining information from multiple sources or applying abstract concepts in novel ways, tasks that remain challenging for current Machine learning approaches. The results highlighted significant gaps between AI capabilities and expert human performance, particularly in areas requiring nuanced judgment or creative synthesis.
Implications and Reception
The release of Humanity's Last Exam generated considerable discussion within the AI research community and beyond. Proponents argued that the benchmark provides a valuable tool for tracking progress toward artificial general intelligence, by setting a high bar that is not easily gamed. Critics, however, noted that benchmarks, no matter how well-designed, can become outdated as models improve, and that performance on such tests does not necessarily translate to real-world competence.
Despite these debates, the benchmark has been widely adopted as a reference point for evaluating frontier models. Several AI companies have reported their scores on the public test set, and the benchmark has been cited in numerous research papers. It has also spurred the development of similar, even more challenging evaluation suites, as the field continues to push the boundaries of what AI can achieve. As of early 2025, no model has come close to passing the benchmark, leaving its name as a testament to the ongoing challenge of creating truly intelligent machines.
Future Directions
The creators of Humanity's Last Exam have indicated plans to update the benchmark periodically, adding new questions and retiring those that become too easy. This iterative approach is intended to maintain its relevance as AI capabilities evolve. There is also ongoing research into automated question generation, which could help scale the benchmark to even larger sizes. The ultimate goal remains to provide a durable, rigorous measure of AI's progress toward expert-level understanding, a goal that continues to shape the development of Generative AI and related fields.