# MMMU-Pro

MMMU-Pro is a professional-grade benchmark for evaluating multimodal AI models on complex, college-level tasks, featuring harder questions and a more rigorous evaluation protocol than the original MMMU benchmark.

MMMU-Pro is a benchmark dataset designed to evaluate the capabilities of multimodal artificial intelligence systems, particularly large multimodal models that process both visual and textual information. It serves as an enhanced and more challenging successor to the original MMMU (Massive Multi-discipline Multimodal Understanding) benchmark, targeting professional-level expertise across various academic and technical disciplines. The benchmark is structured to test a model's ability to integrate and reason over images, diagrams, charts, and associated text in ways that require deep domain knowledge, not just pattern recognition.

The original MMMU benchmark, released in late 2023, consisted of questions drawn from university-level textbooks and exams across six disciplines: art, business, science, humanities, social science, and technology. MMMU-Pro was introduced in 2024 to address limitations in the original, specifically the tendency for some models to achieve high scores by exploiting statistical regularities or memorized answers. MMMU-Pro increases question difficulty, introduces more complex visual elements, and implements a stricter evaluation protocol that includes multiple-choice questions with a larger option set and a more rigorous scoring method, reducing the impact of random guessing.

## Design and Construction

MMMU-Pro was constructed by a team of researchers from multiple institutions, including contributions from academic labs and industry research groups. The dataset builds upon the original MMMU but adds a layer of professional-level scrutiny. Questions are sourced from a wider range of materials, including advanced textbooks, professional certification exams, and research papers. Each question is paired with an image or set of images that are integral to solving the problem, and the text is crafted to require multi-step reasoning.

A key design choice in MMMU-Pro is the expansion of answer choices. While the original MMMU used four options per question, MMMU-Pro typically uses ten options. This change significantly lowers the probability of a model guessing correctly by chance, making the benchmark a more reliable measure of genuine understanding. Additionally, the benchmark includes a subset of questions where the visual information is intentionally misleading or requires careful inspection to extract the relevant details, testing a model's ability to focus on pertinent information over distractors.

## Evaluation Protocol

Evaluation on MMMU-Pro follows a strict protocol to ensure fairness and reproducibility. Models are provided with the question text and the associated image, and they must output one of the ten answer choices. The primary metric is accuracy, but the benchmark also reports performance broken down by discipline and by question type (e.g., diagram interpretation, chart reading, or visual reasoning).

To further reduce the impact of guessing, MMMU-Pro includes a "no-image" condition. In this condition, models are evaluated on the same questions but without the accompanying images. This serves as a control to measure how much a model relies on visual information versus textual priors. A model that performs well on the no-image condition but poorly on the full condition may be overfitting to language patterns, while a model that performs well on both demonstrates robust multimodal integration.

## Performance of Leading Models

As of late 2024, no model has achieved human-level performance on MMMU-Pro. The best-performing systems, which include proprietary models from major AI companies such as [OpenAI](https://www.wikiprompt.org/wiki/openai), [Google DeepMind](https://www.wikiprompt.org/wiki/google-deepmind), and [Anthropic](https://www.wikiprompt.org/wiki/anthropic), typically score in the 50-60% accuracy range on the full benchmark. This is a significant drop from the 70-80% accuracy these same models achieve on the original MMMU, highlighting the increased difficulty.

Open-source models, such as those developed by academic groups and smaller companies, generally score lower, often in the 30-45% range. The gap between proprietary and open-source models is more pronounced on MMMU-Pro than on simpler benchmarks, suggesting that the professional-level reasoning required is a current bottleneck for many [neural network](https://www.wikiprompt.org/wiki/neural-network) architectures. The benchmark has become a standard reference point in the [artificial-intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence) community for tracking progress in multimodal understanding, with researchers regularly reporting results on it in papers and technical reports.

## Impact and Limitations

MMMU-Pro has influenced the development of multimodal models by highlighting specific weaknesses. For instance, many models struggle with questions that require spatial reasoning or the interpretation of complex scientific diagrams, such as those found in [deep-learning](https://www.wikiprompt.org/wiki/deep-learning) research papers. This has spurred work on improving visual encoders and on training techniques that emphasize reasoning over memorization.

However, the benchmark is not without limitations. Critics note that the questions, while difficult, are still multiple-choice, which can allow models to use elimination strategies. Additionally, the dataset is static, meaning that once a model has been trained on it (or on similar data), the results may become inflated. The creators of MMMU-Pro have acknowledged this and have released updated versions with new questions periodically.

Another limitation is the potential for data contamination. Since the benchmark is publicly available, it is possible that some training datasets for large models include MMMU-Pro questions, which would artificially boost scores. Researchers have called for more transparent reporting on training data to mitigate this issue. Despite these concerns, MMMU-Pro remains one of the most rigorous publicly available benchmarks for evaluating professional-level multimodal understanding in [generative AI](https://www.wikiprompt.org/wiki/generative-ai) systems.

## Future Directions

The development of MMMU-Pro is part of a broader trend in AI evaluation towards more challenging and ecologically valid benchmarks. Future iterations may include open-ended questions, interactive tasks, or real-world problem-solving scenarios that go beyond multiple-choice format. The benchmark's emphasis on professional expertise aligns with the growing deployment of AI in fields like medicine, law, and engineering, where errors can have significant consequences.

As models continue to improve, benchmarks like MMMU-Pro will need to evolve to remain relevant. The research community is also exploring complementary evaluation methods, such as human evaluation and adversarial testing, to provide a more complete picture of model capabilities. MMMU-Pro, with its focus on depth and rigor, serves as a valuable tool in this ongoing effort to understand and advance the state of the art in multimodal AI.

---
Source: https://www.wikiprompt.org/wiki/mmmu-pro
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-12T16:26:58.243692+00:00
