# MMMU Benchmark

MMMU is a November 2023 benchmark for evaluating multimodal large language models on college-level, expert reasoning tasks across six disciplines, using 11.5K manually collected questions. It has become a standard evaluation tool for frontier AI systems.

MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark) is a benchmark for evaluating multimodal large language models on tasks that require college-level subject knowledge and deliberate reasoning. Released in November 2023, it was designed to test expert-level perception, knowledge, and reasoning beyond commonsense visual understanding, using questions manually collected from college exams, quizzes, and textbooks.

The benchmark comprises 11.5 thousand multimodal questions spanning six disciplines: Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These cover 30 subjects and 183 subfields, incorporating highly heterogeneous image types such as charts, diagrams, maps, tables, music sheets, and chemical structures. At the time of release, the best-performing models achieved only around 56% accuracy, well below the human expert range of 76–89%, highlighting the benchmark's difficulty.

## Development and Release

MMMU was developed by a team of 22 researchers led by Xiang Yue at Ohio State University and Wenhu Chen at the University of Waterloo. The team manually collected and curated questions to ensure they required genuine college-level knowledge and multi-step reasoning, rather than simple pattern recognition. The benchmark was presented as an oral paper at CVPR 2024, a top computer vision conference, signaling its significance in the field.

## Benchmark Design

The questions in MMMU are designed to be multimodal, meaning they combine text with various image types. This heterogeneity is intentional, as it forces models to integrate visual information from diverse sources, from scientific charts to artistic compositions. The benchmark emphasizes deliberate reasoning, requiring models to apply subject-specific knowledge and logical steps to arrive at answers, rather than relying on superficial cues.

## Impact and Adoption

Since its release, MMMU has become a standard evaluation for frontier multimodal models. Scores on MMMU are routinely reported in the technical reports of systems such as [GPT-4o](https://www.wikiprompt.org/wiki/gpt-4o), [Gemini](https://www.wikiprompt.org/wiki/gemini), and Qwen-VL. The benchmark's difficulty has driven progress in multimodal [artificial intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence), pushing developers to improve models' expert-level perception and reasoning capabilities. Its widespread use reflects a broader trend in [machine learning](https://www.wikiprompt.org/wiki/machine-learning) toward more rigorous, domain-specific evaluation metrics.

## Related Benchmarks and Context

MMMU sits within a landscape of benchmarks aimed at assessing [large language models](https://www.wikiprompt.org/wiki/large-language-model) and multimodal systems. While earlier benchmarks focused on general knowledge or simple visual tasks, MMMU targets college-level expertise and heterogeneous image understanding. This aligns with efforts by organizations like [OpenAI](https://www.wikiprompt.org/wiki/openai), [Anthropic](https://www.wikiprompt.org/wiki/anthropic), and [Google DeepMind](https://www.wikiprompt.org/wiki/google-deepmind) to develop models that can handle complex, real-world problems. The benchmark's emphasis on deliberate reasoning also connects to research in [deep learning](https://www.wikiprompt.org/wiki/deep-learning) and [neural networks](https://www.wikiprompt.org/wiki/neural-network), particularly in areas like [transformers](https://www.wikiprompt.org/wiki/transformer) and [multi-head attention](https://www.wikiprompt.org/wiki/multi-head-attention).

## Limitations and Future Directions

Despite its success, MMMU has limitations. The manual collection process is labor-intensive, and the benchmark's static nature may lead to saturation as models improve. Researchers have noted that high scores on MMMU do not necessarily translate to robust performance in dynamic, real-world settings. Future benchmarks may need to incorporate more adaptive or generative evaluation methods, building on MMMU's foundation to address these challenges.

## References

- Yue, X., et al. (2023). MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark. arXiv preprint.
- CVPR 2024 oral presentation.

---
Source: https://www.wikiprompt.org/wiki/mmmu-benchmark
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-13T03:50:16.527228+00:00
