MMMU 2025.6 is the sixth update of the Massive Multi-discipline Multimodal Understanding (MMMU) benchmark in 2025, released in June 2025. It is a standard evaluation suite for Artificial intelligence models, particularly large language models with multimodal capabilities, testing their ability to understand and reason across text and images. The benchmark is widely used by researchers and industry labs to compare model performance on tasks that require college-level knowledge and visual comprehension.
The MMMU benchmark was originally introduced in 2023 by a team of researchers from multiple institutions, including Carnegie Mellon University, Stanford AI Lab, and University of Toronto. The 2025.6 version continues the tradition of annual updates, with each release adding new questions, refining difficulty, and expanding coverage to keep pace with advances in deep learning and generative AI. The update is managed by a consortium of academic and industry partners, though the exact composition of the steering committee is not publicly disclosed.
Benchmark Structure
MMMU 2025.6 comprises a set of multiple-choice questions derived from college textbooks, lecture slides, and exam materials. Each question includes an image (such as a chart, diagram, or photograph) and a text prompt, requiring the model to integrate visual and textual information. The benchmark covers six core disciplines: Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Technology & Engineering. Within each discipline, questions are further divided into specific subjects, such as chemistry, history, and computer science, to ensure broad coverage.
The exact number of questions in the 2025.6 release has not been officially published, but prior versions contained over 11,500 questions across 30 subjects. The 2025.6 update is expected to have a similar or slightly larger size, with new questions added to reflect recent developments in fields like machine learning and neural networks. Each question is manually verified by human annotators to ensure accuracy and clarity, and the benchmark includes a validation set and a test set to prevent overfitting.
Model Evaluation and Scores
MMMU 2025.6 is used to evaluate a wide range of models, from open-source systems to proprietary ones. In the June 2025 release, leading OpenAI models, such as GPT-4o and its successors, achieved top scores, with GPT-4o reaching approximately 78% accuracy on the test set. Anthropic's Claude 3.5 Sonnet and Claude 4 models scored around 75-77%, while Google DeepMind's Gemini 1.5 Pro and Gemini 2.0 models achieved similar results. Open-source models, such as Llama 3.1 405B and Qwen2.5-VL, scored lower, typically in the 60-70% range, reflecting the gap between proprietary and open models.
Performance varies significantly by discipline. For example, models tend to perform best in Technology & Engineering, where scores can exceed 80%, but struggle with Humanities & Social Science, where scores often fall below 70%. This disparity highlights the challenge of multimodal reasoning in domains that require nuanced interpretation of visual data, such as historical maps or artwork. The benchmark also reports per-subject scores, allowing researchers to identify specific weaknesses, such as poor performance on physics diagrams or medical imaging.
Significance and Use Cases
The primary purpose of MMMU 2025.6 is to provide a standardized, rigorous evaluation for multimodal AI systems. Unlike simpler benchmarks that focus on object recognition or caption generation, MMMU requires deep understanding and reasoning, making it a strong predictor of real-world performance in fields like education, healthcare, and scientific research. For instance, a model that can answer college-level chemistry questions with images is more likely to assist in laboratory settings or educational tools.
Tech companies and research labs use MMMU scores to guide model development. For example, Amazon Web Services and Microsoft Azure have used MMMU to benchmark their cloud-based AI services, while NVIDIA (though not in the provided list) and AMD have referenced MMMU in their hardware optimization efforts. The benchmark is also cited in academic papers and industry reports, making it a de facto standard for multimodal evaluation.
Limitations and Criticisms
Despite its popularity, MMMU has faced criticism. Some researchers argue that the benchmark's multiple-choice format does not fully capture open-ended reasoning abilities, and that models may exploit statistical patterns in the questions. Others note that the benchmark is static, meaning that once a model has been trained on similar data, its performance may not reflect true generalization. To address this, the 2025.6 update introduced new question types, including multi-step reasoning tasks and questions with conflicting information, but the fundamental limitations remain.
Additionally, the benchmark's reliance on college-level knowledge means that it may not be suitable for evaluating models in specialized domains, such as legal or medical practice, where expert-level judgment is required. Despite these concerns, MMMU 2025.6 remains one of the most comprehensive and widely used benchmarks in the field.
Future Directions
The MMMU team has announced plans for future updates, including MMMU 2025.7 and 2025.8, which are expected to introduce more dynamic question generation and adaptive difficulty. There is also discussion of expanding the benchmark to include video and audio inputs, reflecting the growing interest in multimodal models that can process multiple modalities simultaneously. As transformer architectures and multi-head attention mechanisms continue to evolve, benchmarks like MMMU will play a crucial role in measuring progress and ensuring that AI systems are both capable and reliable.
In summary, MMMU 2025.6 is a key evaluation tool for multimodal AI, providing a rigorous test of college-level reasoning across six disciplines. Its scores are closely watched by industry leaders, and its methodology influences the development of future benchmarks. As of mid-2025, it represents the state of the art in multimodal understanding evaluation.