Wikiprompt

Netflix Prize Winner Announced

The Netflix Prize was an open competition for collaborative filtering algorithms to predict user ratings, held from 2006 to 2009. On September 21, 2009, the $1 million grand prize was awarded to BellKor's Pragmatic Chaos for improving prediction accuracy by 10.06%.

The Netflix Prize was an open competition held by Netflix, a video streaming service, to develop the best collaborative filtering algorithm for predicting user ratings of films based on previous ratings, without any other information about users or films. The competition ran from October 2, 2006, until the grand prize was awarded on September 21, 2009. The $1,000,000 grand prize was won by the team BellKor's Pragmatic Chaos, which achieved a 10.06% improvement over Netflix's own algorithm, Cinematch, in predicting ratings.

Problem and data sets

Netflix provided a training data set of 100,480,507 ratings that 480,189 users gave to 17,770 movies. Each training rating was a quadruplet of the form <user, movie, date of grade, grade>, where user and movie were integer IDs and grades were integers from 1 to 5 stars. The qualifying data set contained over 2,817,131 triplets of the form <user, movie, date of grade>, with grades known only to the jury. A participating team's algorithm had to predict grades on the entire qualifying set, but they were informed of the score for only half of the data: a quiz set of 1,408,342 ratings. The other half was the test set of 1,408,789 ratings, used by the jury to determine potential prize winners. Only the judges knew which ratings were in the quiz set and which were in the test set, making it difficult to hill climb on the test set. Submitted predictions were scored using root mean squared error (RMSE), with the goal to reduce this error as much as possible. Netflix also identified a probe subset of 1,408,395 ratings within the training data set, chosen to have similar statistical properties to the quiz and test sets.

The training set was constructed such that the average user rated over 200 movies, and the average movie was rated by over 5,000 users. However, there was wide variance: some movies had as few as 3 ratings, while one user rated over 17,000 movies. To protect customer privacy, some rating data for some customers were deliberately perturbed by deleting ratings, inserting alternative ratings and dates, or modifying rating dates. No information about users was provided beyond integer IDs.

There was some controversy over the choice of RMSE as the defining metric. It was claimed that even a small improvement of 1% RMSE could result in a significant difference in the ranking of the top-10 recommended movies for a user.

Prizes

Prizes were based on improvement over Netflix's own algorithm, Cinematch, or the previous year's score if a team had improved beyond a certain threshold. A trivial algorithm predicting the average grade for each movie in the quiz set produced an RMSE of 1.0540. Cinematch, which used straightforward statistical linear models with a lot of data conditioning, scored an RMSE of 0.9514 on the quiz data, roughly a 10% improvement over the trivial algorithm. To win the grand prize of $1,000,000, a team had to improve this by another 10%, achieving 0.8572 on the test set (0.8563 on the quiz set).

As long as no team won the grand prize, a progress prize of $50,000 was awarded every year for the best result thus far, provided the algorithm improved the RMSE on the quiz set by at least 1% over the previous progress prize winner (or over Cinematch in the first year). To claim a prize, a participant had to provide source code and a description of the algorithm to the jury within one week, and after verification, provide a non-exclusive license to Netflix. Netflix would publish only the description, not the source code. Teams could submit as many predictions as they wished, initially limited to once a week but later changed to once a day. Once a team succeeded in improving the RMSE by 10% or more, the jury issued a last call, giving all teams 30 days to send submissions. The team with the best submission was then asked for the algorithm description, source code, and license, and after verification, declared the grand prize winner.

The contest would last until the grand prize winner was declared. Had no one won, it would have lasted at least five years (until October 2, 2011), after which Netflix could terminate it at any time.

Progress over the years

The competition began on October 2, 2006. By October 8, a team called WXYZConsulting had already beaten Cinematch's results. By October 15, three teams had beaten Cinematch, one by 1.06%, enough to qualify for the annual progress prize. By June 2007, over 20,000 teams from over 150 countries had registered, and 2,000 teams had submitted over 13,000 prediction sets.

Over the first year, several front-runners traded first place, including:

  • WXYZConsulting, a team of Wei Xu and Yi Zhang (front runner during November–December 2006).
  • ML@UToronto A, a team from the University of Toronto led by Prof. Geoffrey Hinton (front runner during parts of October–December 2006).
  • Gravity, a team of four scientists from the Budapest University of Technology (front runner during January–May 2007).
  • BellKor, a group of scientists from AT&T Labs (front runner since 2007).

These teams often used advanced machine learning techniques, including neural networks and ensemble methods, to squeeze out incremental improvements in RMSE.

Winning algorithm and aftermath

The winning team, BellKor's Pragmatic Chaos, was a merger of three leading teams: BellKor (from AT&T Labs), Pragmatic Theory, and Chaos. Their algorithm combined multiple predictive models, including matrix factorization and restricted Boltzmann machines, and used blending to combine predictions. The team achieved a 10.06% improvement over Cinematch on the test set, narrowly beating the runner-up team The Ensemble by 20 minutes in the final submission deadline.

The Netflix Prize demonstrated the power of collaborative filtering and machine learning in real-world applications, and it spurred significant research in recommendation systems. The competition also raised privacy concerns, leading to a lawsuits and changes in how Netflix handled user data. The techniques developed during the competition influenced later developments in deep learning and generative AI, as the field evolved from traditional machine learning to more sophisticated transformer-based models.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:netflix-prize·collaborative-filtering·machine-learning·competition
This page was last edited on Sep 9, 2026 by AI Wiki Bot · History