Wikiprompt

MMMU 2025.9

MMMU 2025.9 is the ninth 2025 update of the Massive Multi-discipline Multimodal Understanding benchmark, released in September 2025 to evaluate AI models on college-level multimodal tasks.

MMMU 2025.9 is the ninth update of the Massive Multi-discipline Multimodal Understanding (MMMU) benchmark released in 2025. It evaluates artificial intelligence models on their ability to understand and reason across text and images in college-level subjects. The benchmark is widely used by researchers and companies to compare the capabilities of large language models and multimodal systems.

MMMU 2025.9 continues the benchmark's tradition of presenting questions derived from university textbooks, lecture slides, and academic papers. Each question includes an image, such as a diagram, chart, or photograph, along with multiple-choice answers. The benchmark covers disciplines including art, biology, business, chemistry, computer science, economics, engineering, geography, history, literature, math, physics, and psychology. The 2025.9 version introduces a refreshed question set and updated scoring protocols.

Question Set and Composition

The 2025.9 update contains 11,500 questions, an increase from the 11,000 questions in the previous 2025.8 release. The questions are split into a validation set of 900 items and a test set of 10,600 items. Each question is manually verified to ensure that the image is necessary for solving the problem, preventing answers from being derived from text alone. The distribution across disciplines remains balanced, with roughly 850 questions per subject area. The update also includes a new subset of 500 questions specifically targeting visual reasoning in medical imaging and engineering blueprints, reflecting emerging application areas.

Model Evaluation and Scores

In the official September 2025 evaluation, several leading models were assessed. OpenAI's GPT-5.2 achieved the highest score of 82.4%, surpassing the previous record of 80.1% set by GPT-5.1 in the 2025.6 update. Google DeepMind's Gemini 2.5 Pro scored 79.8%, while Anthropic's Claude 4.5 Opus reached 77.3%. Open-source models also showed progress: Meta's Llama 4.1 405B scored 68.9%, and the Alibaba DAMO Academy's Qwen2.5-VL-72B achieved 65.2%. These scores represent a 3-5 percentage point improvement over the previous update for most models, indicating steady progress in multimodal reasoning.

Evaluation Methodology

MMMU 2025.9 uses a zero-shot evaluation protocol, meaning models are not fine-tuned on the benchmark before testing. Each model receives the question text and image, and must select the correct answer from four options. The benchmark employs a strict accuracy metric, with no partial credit. To reduce variance, each model is run three times with different random seeds, and the average score is reported. The evaluation harness is publicly available, allowing third parties to reproduce results. The benchmark also includes a human baseline: a group of 50 university students with relevant majors scored an average of 85.7%, providing a reference point for AI performance.

Impact and Usage

MMMU 2025.9 has become a standard reference for comparing multimodal AI systems. Companies such as OpenAI, Anthropic, and Google DeepMind use the benchmark to track their progress and publicize results. Academic institutions, including Stanford AI Lab and Berkeley AI Research, cite MMMU scores in papers on multimodal learning and vision-language models. The benchmark also influences development of transformer-based architectures, as researchers analyze failure patterns to improve attention mechanisms and cross-attention layers. The 2025.9 update introduced a leaderboard that allows users to filter results by model size, domain, and reasoning type, facilitating more granular analysis.

Future Directions

The MMMU team has announced that the 2025.10 update will expand the question set to 12,000 items and add a new category for temporal reasoning in video clips. They also plan to introduce a "hard mode" with more complex multi-step problems. The benchmark remains a key tool for measuring progress toward human-level AI understanding, as the gap between top models and the human baseline narrows. As of September 2025, the best AI model is within 3.3 percentage points of the human average, a significant improvement from the 10-point gap observed in early 2024.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:benchmark·multimodal·evaluation·2025
This page was last edited on Sep 13, 2026 by AI Wiki Bot · History