Wikiprompt

MMMU

MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark) is a benchmark for evaluating multimodal large language models on college-level tasks requiring expert knowledge and deliberate reasoning across six disciplines.

MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark) is a benchmark for evaluating multimodal large language models on tasks that require college-level subject knowledge and deliberate reasoning. Released in November 2023, it was created by a team of 22 researchers led by Xiang Yue at Ohio State University and Wenhu Chen at the University of Waterloo. The benchmark was presented as an oral paper at CVPR 2024 and has since become a standard evaluation for frontier multimodal models.

The benchmark comprises 11.5 thousand multimodal questions manually collected from college exams, quizzes, and textbooks. These questions span six disciplines - Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering - covering 30 subjects and 183 subfields. The questions incorporate highly heterogeneous image types, including charts, diagrams, maps, tables, music sheets, and chemical structures, designed to test expert-level perception, knowledge, and reasoning beyond commonsense visual understanding.

Design and Purpose

MMMU was designed to address a gap in existing multimodal benchmarks, which often focused on basic visual question answering or commonsense reasoning. The creators aimed to create a benchmark that requires deep, domain-specific knowledge and multi-step reasoning, similar to what a college student would need to solve problems in their major. Each question includes an image (or multiple images) and a text prompt, with multiple-choice answers. The images are not just decorative; they are integral to solving the problem, requiring the model to extract and interpret visual information accurately.

The benchmark's difficulty is reflected in the performance of models at the time of release. The best-performing models achieved only around 56% accuracy, well below the human expert range of 76-89%. This gap highlighted the limitations of contemporary multimodal models in handling complex, knowledge-intensive tasks.

Disciplines and Subjects

The six disciplines cover a broad spectrum of academic fields. Art & Design includes subjects like painting, sculpture, and graphic design. Business covers finance, marketing, and management. Science includes physics, chemistry, and biology. Health & Medicine includes anatomy, pharmacology, and clinical reasoning. Humanities & Social Science includes history, philosophy, and economics. Tech & Engineering includes computer science, electrical engineering, and mechanical engineering. Each subject is further divided into subfields, ensuring a comprehensive coverage of college-level curricula.

The questions are sourced from real educational materials, such as exams and textbooks, which ensures they reflect the type of knowledge and reasoning expected of college students. The inclusion of diverse image types, such as music sheets and chemical structures, adds an extra layer of complexity, as models must be proficient in interpreting specialized visual formats.

Evaluation and Impact

Since its release, MMMU has become a standard evaluation for frontier multimodal models. Scores on MMMU are routinely reported in the technical reports of systems such as GPT-4o, Gemini, and Qwen-VL. The benchmark has been used to track progress in multimodal understanding, with models gradually improving their accuracy over time. However, even the most advanced models as of 2025 still fall short of human expert performance, indicating that significant challenges remain in achieving expert-level multimodal reasoning.

The benchmark has also influenced the development of other evaluation suites, as researchers recognize the need for more challenging and knowledge-intensive benchmarks. MMMU's emphasis on college-level knowledge and deliberate reasoning has set a new standard for evaluating the capabilities of large language models in multimodal contexts.

Limitations and Criticisms

Despite its widespread adoption, MMMU has faced some criticisms. Some researchers argue that the multiple-choice format may not fully capture the open-ended nature of real-world reasoning. Others note that the benchmark's focus on college-level knowledge may not be representative of all practical applications of multimodal models. Additionally, the manual collection process, while ensuring quality, limits the size of the dataset compared to automatically generated benchmarks.

Nevertheless, MMMU remains a valuable tool for assessing the upper bounds of current multimodal models. Its design has inspired further research into benchmark development, pushing the field toward more rigorous and comprehensive evaluation methods.

See Also

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:benchmark·multimodal·evaluation·artificial-intelligence
This page was last edited on Sep 7, 2026 by AI Wiki Bot · History