The Netflix Prize was an open competition held by Netflix, a video streaming service, to develop the best collaborative filtering algorithm for predicting user ratings of films. The competition ran from October 2, 2006, to September 21, 2009, and was open to anyone not connected with Netflix (including current and former employees and their close relatives) and not a resident of certain blocked countries such as Cuba or North Korea. The grand prize of US$1,000,000 was awarded to the BellKor's Pragmatic Chaos team, which achieved a 10.06% improvement over Netflix's own algorithm, Cinematch, in predicting ratings.
Problem and Data Sets
Netflix provided a training data set containing 100,480,507 ratings that 480,189 users gave to 17,770 movies. Each training rating was a quadruplet of the form <user, movie, date of grade, grade>, where user and movie were integer IDs and grades were integers from 1 to 5 stars. The qualifying data set contained over 2,817,131 triplets of the form <user, movie, date of grade>, with grades known only to the jury. A participating team's algorithm had to predict grades on the entire qualifying set, but teams were informed of the score for only half of the data: a quiz set of 1,408,342 ratings. The other half, the test set of 1,408,789 ratings, was used by the jury to determine potential prize winners. Only the judges knew which ratings were in the quiz set and which were in the test set, an arrangement intended to make it difficult to hill climb on the test set.
Submitted predictions were scored against the true grades using root mean squared error (RMSE), and the goal was to reduce this error as much as possible. While actual grades were integers in the range 1 to 5, submitted predictions did not need to be. Netflix also identified a probe subset of 1,408,395 ratings within the training data set. The probe, quiz, and test data sets were chosen to have similar statistical properties. In summary, the data used in the Netflix Prize was as follows: training set (99,072,112 ratings not including the probe set; 100,480,507 including the probe set), probe set (1,408,395 ratings), and qualifying set (2,817,131 ratings) consisting of the test set (1,408,789 ratings) and quiz set (1,408,342 ratings).
For each movie, the title and year of release were provided in a separate dataset, but no information was provided about users. To protect customer privacy, some rating data for some customers in the training and qualifying sets were deliberately perturbed by deleting ratings, inserting alternative ratings and dates, and modifying rating dates. The training set was constructed such that the average user rated over 200 movies, and the average movie was rated by over 5,000 users, but there was wide variance: some movies had as few as 3 ratings, while one user rated over 17,000 movies. There was some controversy regarding the choice of RMSE as the defining metric, as it was claimed that even a 1% improvement in RMSE could result in a significant difference in the ranking of the top-10 most recommended movies for a user.
Prizes
Prizes were based on improvement over Netflix's own algorithm, called Cinematch, or over the previous year's score if a team had made improvement beyond a certain threshold. A trivial algorithm that predicted for each movie in the quiz set its average grade from the training data produced an RMSE of 1.0540. Cinematch used straightforward statistical linear models with a lot of data conditioning, and its performance had plateaued by 2006. Using only the training data, Cinematch scored an RMSE of 0.9514 on the quiz data, roughly a 10% improvement over the trivial algorithm, and had a similar performance on the test set at 0.9525. To win the grand prize of $1,000,000, a participating team had to improve this by another 10%, achieving 0.8572 on the test set, which corresponded to an RMSE of 0.8563 on the quiz set.
As long as no team won the grand prize, a progress prize of $50,000 was awarded every year for the best result thus far. To win this prize, an algorithm had to improve the RMSE on the quiz set by at least 1% over the previous progress prize winner (or over Cinematch in the first year). If no submission succeeded, the progress prize was not awarded for that year. To win a progress or grand prize, a participant had to provide source code and a description of the algorithm to the jury within one week after being contacted, and following verification, the winner also had to provide a non-exclusive license to Netflix. Netflix would publish only the description, not the source code, of the system. A team could choose not to claim a prize to keep their algorithm and source code secret. The jury kept their predictions secret from other participants. Teams could send as many attempts to predict grades as they wished, with submissions originally limited to once a week but quickly modified to once a day. A team's best submission so far counted as their current submission.
Once a team succeeded in improving the RMSE by 10% or more, the jury would issue a last call, giving all teams 30 days to send their submissions. Only then was the team with the best submission asked for the algorithm description, source code, and non-exclusive license, and after successful verification, declared the grand prize winner. The contest would last until the grand prize winner was declared, but had no one received the grand prize, it would have lasted for at least five years (until October 2, 2011), after which it could have been terminated at any time at Netflix's sole discretion.
Progress Over the Years
The competition began on October 2, 2006. By October 8, a team called WXYZConsulting had already beaten Cinematch's results. By October 15, there were three teams that had beaten Cinematch, one of them by 1.06%, enough to qualify for the annual progress prize. By June 2007, over 20,000 teams had registered for the competition from over 150 countries, and 2,000 teams had submitted over 13,000 prediction sets.
Over the first year of the competition, a handful of front-runners traded first place. The more prominent ones included WXYZConsulting, a team of Wei Xu and Yi Zhang, which was a front runner during November and December 2006; ML@UToronto A, a team from the University of Toronto led by Prof. Geoffrey Hinton, a front runner during parts of October through December 2006; Gravity, a team of four scientists from the Budapest University of Technology, a front runner during January through May 2007; and BellKor, a group of scientists from AT&T Labs, a front runner since mid-2007. The competition saw the application of various machine learning techniques, including neural networks and deep learning approaches, which were later influential in the broader field of artificial intelligence.
Legacy and Impact
The Netflix Prize demonstrated the effectiveness of ensemble methods and matrix factorization techniques in collaborative filtering. It spurred significant academic and industrial research in recommendation systems, influencing later developments in generative AI and large-scale language models. The competition also raised awareness of privacy issues in data sharing, as researchers later showed that the anonymized data could be re-identified, leading to a 2010 lawsuit and subsequent changes in how companies handle user data. The prize is often cited as a landmark event in the history of machine learning and crowdsourced problem solving, and it inspired similar competitions in other domains.
Conclusion
The Netflix Prize concluded on September 21, 2009, when the grand prize was awarded to BellKor's Pragmatic Chaos, a team that combined the efforts of BellKor, Pragmatic Theory, and BigChaos. The winning algorithm achieved an RMSE of 0.8567 on the test set, a 10.06% improvement over Cinematch. The competition not only advanced the state of the art in recommendation algorithms but also highlighted the potential of open innovation and the importance of data quality and privacy in the digital age. Its influence persists in modern recommendation systems used by streaming services and e-commerce platforms.