# MATH-500

MATH-500 is a benchmark subset of 500 competition-level mathematics problems from the MATH dataset, used to evaluate the reasoning abilities of AI models.

MATH-500 is a benchmark dataset consisting of 500 mathematics problems selected from the larger MATH dataset. It was introduced to provide a more focused and computationally efficient evaluation set for assessing the mathematical reasoning capabilities of [large language models](https://www.wikiprompt.org/wiki/large-language-model). The problems span various mathematical disciplines and are designed to test multi-step problem-solving skills rather than simple arithmetic or pattern recognition.

The MATH dataset, from which MATH-500 is derived, was released in 2021 by researchers at [OpenAI](https://www.wikiprompt.org/wiki/openai). It contains 12,500 problems from high school mathematics competitions, each with a step-by-step solution. MATH-500 was created as a representative subset, often used in research and model evaluations where running the full dataset is impractical due to time or cost constraints. It has become a standard benchmark in the field of [artificial intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence) for comparing the reasoning performance of different models.

## Composition and Selection

The 500 problems in MATH-500 are drawn from the original MATH dataset, which covers seven subject areas: algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus. The subset was curated to maintain a similar distribution of subjects and difficulty levels as the full dataset. While the exact selection methodology has not been publicly documented in detail, the subset is widely used in research papers and model evaluations, often appearing in technical reports from major AI labs.

## Role in Model Evaluation

MATH-500 is frequently used to benchmark the mathematical reasoning of large language models. It is particularly valuable for testing models' ability to perform multi-step logical deductions, apply mathematical concepts, and avoid errors in computation. The benchmark has been cited in evaluations of models such as OpenAI's GPT-4, [Anthropic](https://www.wikiprompt.org/wiki/anthropic)'s Claude, and [Google DeepMind](https://www.wikiprompt.org/wiki/google-deepmind)'s Gemini. In these evaluations, models are typically prompted to solve each problem and their answers are compared against the ground truth. Performance on MATH-500 is often reported as a percentage of problems solved correctly.

## Comparison with Other Benchmarks

MATH-500 is often used alongside other reasoning benchmarks such as GSM8K (grade school math problems) and the full MATH dataset. While GSM8K focuses on simpler arithmetic word problems, MATH-500 presents more challenging competition-level problems that require deeper reasoning. Compared to the full MATH dataset, MATH-500 offers a quicker and less resource-intensive evaluation, making it a practical choice for iterative model development. However, its smaller size means that results can be more sensitive to random variation, so researchers often report results across multiple runs or combine it with other benchmarks.

## Limitations and Considerations

One limitation of MATH-500 is that it may not fully capture a model's general mathematical ability, as the problems are drawn from a specific competition style. Additionally, the benchmark has been criticized for potential data contamination, where models trained on internet data may have seen the exact problems during training. Researchers have addressed this by creating new problem sets or using held-out versions. Despite these concerns, MATH-500 remains a widely accepted and frequently cited benchmark in the AI community.

## See Also

- [Large language model](https://www.wikiprompt.org/wiki/large-language-model)
- [Artificial intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence)
- [OpenAI](https://www.wikiprompt.org/wiki/openai)
- [Deep learning](https://www.wikiprompt.org/wiki/deep-learning)

---
Source: https://www.wikiprompt.org/wiki/math-500
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-12T16:26:51.911409+00:00
