The MNIST database (Modified National Institute of Standards and Technology database) is a large database of handwritten digits that is commonly used for training various image processing systems. It is also widely used for training and testing in the field of Machine learning. The database was created by "re-mixing" the samples from NIST's original datasets, as the creators felt that the original training dataset, taken from American Census Bureau employees, and the testing dataset, taken from American high school students, were not well-suited for machine learning experiments. The black and white images from NIST were normalized to fit into a 28x28 pixel bounding box and anti-aliased, which introduced grayscale levels.
The MNIST database contains 60,000 training images and 10,000 testing images. Half of the training set and half of the test set were taken from NIST's training dataset, while the other half of each came from NIST's testing dataset. The original creators maintain a list of methods tested on it; in their original paper, they used a support-vector machine to achieve an error rate of 0.8%. The original MNIST dataset contains at least four wrong labels.
History
USPS database
In 1988, a dataset of digits from the US Postal Service was constructed. It contained 16×16 grayscale images digitized from handwritten zip codes on U.S. mail passing through the Buffalo, New York post office. The training set had 7,291 images, and the test set had 2,007, totaling 9,298. Both sets contained ambiguous, unclassifiable, and misclassified data. This dataset was used to train and benchmark the 1989 LeNet. The task was difficult; on the test set, two humans made errors at an average rate of 2.5%.
Special Database
In the late 1980s, the Census Bureau was interested in automatic digitization of handwritten census forms and enlisted the Image Recognition Group (IRG) at NIST to evaluate OCR systems. Several years of work resulted in several "Special Databases" and benchmarks. Of particular importance to MNIST are Special Database 1 (SD-1), released in May 1990, Special Database 3 (SD-3), released in February 1992, and Special Database 7 (SD-7), or NIST Test Data 1 (TD-1), released in April 1992. They were released on ISO-9660 CD-ROMs. The data was obtained by asking people to write on "Handwriting Sample Forms" (HSFs), then digitizing the forms and segmenting out the alphanumerical characters. Each writer wrote a single HSF.
Each HSF contains multiple entry fields, including name and date entries, a city/state field, 28 digit fields, one upper-case field, one lower-case field, and an unconstrained Constitution text paragraph. Each HSF was scanned at a resolution of 300 dots per inch (11.8 dots per millimeter).
SD-1 and SD-3 were constructed from the same set of HSFs by 2,100 out of 3,400 permanent census field workers as part of the 1990 United States census. SD-1 contained the segmented data entry fields, but not the segmented alphanumericals. SD-3 contained binary 128×128 images digitized from segmented alphanumericals, with 223,125 digits, 44,951 upper-case letters, and 45,313 lower-case letters.
SD-7 (or TD-1) was the test set, containing 58,646 128×128 binary images written by 500 high school students in Bethesda, Maryland, described as "math and science students in a high school as a short exercise during class". Each image was accompanied by a unique integer ID for the writer's identity. SD-7 was released without labels on CD-ROMs, with labels later released on floppy drives. It did not contain the HSFs. The human error rate on SD-7 was 1.5%.
SD-3 was much cleaner and easier to recognize than SD-7. The European crossed seven (7) was far more abundant in SD-7 than in SD-3. It was suspected that SD-3 was produced by more motivated people than SD-7, and that the character segmenter for SD-3, an older design, failed more often, filtering out harder instances. Machine learning systems trained and validated on SD-3 suffered significant drops in performance on SD-7, from an error rate of less than 1% to about 10%.
In 1992, NIST and the Census Bureau sponsored a competition and conference to determine the state of the art. Teams were given SD-3 as the training set before March 23, SD-7 as the test set before April 13, and would submit systems for classifying SD-7 before April 27. A total of 45 algorithms were submitted from 26 companies in 7 countries. On May 27 and 28, all parties convened in Gaithersburg, Maryland at the First Census OCR Systems Conference, with observers from FBI, IRS, and USPS. The winning entry did not use SD-3 for training but a much larger proprietary training set, avoiding the distribution shift. Among the 25 entries that used SD-3, the winner was a nearest-neighbor classifier using a handcrafted metric invariant to Euclidean transforms.
SD-19 was published in 1995 as a compilation of SD-1, SD-3, SD-7, and further data. It contained 814,255 binary images of alphanumericals and binary images of 4,169 HSFs, including the 500 HSFs used to generate SD-7. It was updated in 2016.
Construction of MNIST
The MNIST was constructed sometime before summer 1994 by mixing 128x128 binary images from SD-3 and SD-7. Specifically, the creators first took all images from SD-7 and divided them into a training set and a test set, each from 250 writers, resulting in nearly 30,000 images in each set. They then added more images from SD-3 until each set contained exactly 60,000 images.
Each image was size-normalized to fit in a 20x20 pixel box while preserving its aspect ratio, and anti-aliased to grayscale. Then it was placed into a 28x28 image by translating it until the center of mass of the pixels was in the center of the image. The details of the downsampling procedure were later reconstructed.
The training set and the test set both originally had 60,000 samples, but 50,000 of the test set samples were discarded, leaving only the samples indexed 24,476 to 34,475, giving just 10,000 samples in the test set.
Further versions
In 2019, the full 60,000 test set from MNIST was restored to construct the QMNIST, which has 60,000 images in the training set and 60,000 in the test set.
Extended MNIST (EMNIST) is a newer dataset developed and released by NIST as the successor to MNIST, released in 2017. While MNIST included only handwritten digits, EMNIST was constructed from all the images in SD-19, including letters and digits, and provides a more challenging benchmark for Deep learning and Artificial intelligence research.
Legacy
MNIST has become a fundamental benchmark in Machine learning education and research, often used as a first test for new algorithms and architectures. Its simplicity and small size make it ideal for teaching concepts in Neural network training and evaluation. The dataset's widespread adoption has influenced the development of subsequent benchmarks and contributed to the growth of the field, including advances in Deep learning and Generative AI.