The Netflix Prize was an open competition held by Netflix, a video streaming service, to develop the best collaborative filtering algorithm for predicting user ratings of films. The competition ran from October 2, 2006, to September 21, 2009, and was open to anyone not connected with Netflix (including current and former employees and their close relatives) and not a resident of certain blocked countries such as Cuba or North Korea. The grand prize of US$1,000,000 was awarded to BellKor's Pragmatic Chaos, a team that combined the efforts of BellKor, Pragmatic Theory, and BigChaos.
The competition used a training data set of over 100 million ratings that 480,189 users gave to 17,770 movies. The qualifying data set contained over 2.8 million ratings. The test set consisted of over 1.4 million ratings. The goal was to predict these ratings with a root mean squared error (RMSE) that was at least 10% better than Netflix's own Cinematch algorithm. The winning algorithm achieved an RMSE of 0.8567 on the test set, a 10.06% improvement over Cinematch.
The prize is often cited as a landmark event in the history of machine learning and crowdsourced problem solving. It demonstrated the effectiveness of ensemble methods and matrix factorization techniques in collaborative filtering. It also raised awareness of privacy issues in data sharing, as researchers later showed that the anonymized data could be re-identified, leading to a 2010 lawsuit and subsequent changes in how companies handle user data. The competition inspired similar competitions in other domains and its influence persists in modern recommendation systems used by streaming services and e-commerce platforms.