The 2020 research paper "Denoising Diffusion Probabilistic Models," authored by Jonathan Ho, Ajay Jain, and Pieter Abbeel, introduced a practical and high-quality method for generative modeling using diffusion processes. Building on earlier work in non-equilibrium thermodynamics from 2015, the paper proposed a specific framework called the Denoising Diffusion Probabilistic Model (DDPM). In machine learning, diffusion models are a class of latent variable generative models that learn to produce new data by reversing a gradual noising process. They have become central to modern generative AI, especially for image generation, alongside other approaches in deep learning.
The core idea of a diffusion model is two-part: a forward diffusion process that gradually adds noise to data (until it resembles pure Gaussian noise), and a reverse sampling process that learns to remove noise, recovering the original data distribution. A trained model can generate new samples by starting from random noise and iteratively denoising it. The 2020 paper proposed DDPM, which improved upon earlier methods by introducing a simplified variational bound and specific parameterization.
Concepts and formalism
Diffusion models are a class of latent variable models that model data as generated by a diffusion process - a random walk with drift through the space of possible data points. The forward process is a Markov chain that adds Gaussian noise over T steps, often using a schedule of constants β_t. The reverse process is also a Markov chain, parameterized by a neural network, that learns to predict the noise added at each step. The model is trained using variational inference with a simple mean-squared error loss.
A key innovation of the DDPM paper is showing that the training objective simplifies to a weighted mean-squared error term matching the model's predicted noise to actual noise. The paper also established that a simpler objective without additional weighting terms works well in practice. The reverse diffusion model, often called the "backbone," can be any network, typically a U-Net or a transformer. For images, the backbone is usually a U-Net variant applied iteratively to denoise image samples.
How DDPM Works
The forward diffusion process is defined by a sequence of noise levels β_1,...,β_T, with α_t = 1 − β_t and ᾱ_t as the cumulative product. At each step t, Gaussian noise is added to the image. A critical insight is that after t steps, the noisy image can be expressed directly as a combination of the original image and Gaussian noise, allowing efficient training at random timesteps. The reverse process then learns to estimate the noise that was added, enabling recovery of the original data.
In the DDPM, training involves randomly selecting timesteps and minimizing a loss that compares the noise-prediction network output with actual Gaussian noise. During sampling, the model starts with pure Gaussian noise and iteratively applies the learned reverse process. The paper also noted that simpler uniform weighting over timesteps works the best, and that early stages are slower in reverse.
Impact
The method became foundational for many subsequent diffusion models. It demonstrated high-quality image generation on datasets such as CIFAR-10 and LSUN, challenging then-dominant generative adversarial networks. Since then, diffusion models have been widely adopted in Generative AI and Machine learning, leading to commercial systems like DALL-E and Stable Diffusion, which combine diffusion models with text encoders and Cross-Attention for text-conditioned generation. The backbone is often a transformer and U-net hybrid, with some models using a Transformer (architecture) as the backbone.
Legacy and Impact
The 2020 paper catalyzed a shift in Generative AI development. Prior to diffusion models, generative-adversarial-networks ruled image generation, but DDPM surpassed them in sample quality and training stability. As of 2024, diffusion models are mainly used for tasks like image denoising, inpainting, super-resolution, image generation, and video generation. They have also been applied to other domains, including text generation, sound generation, and reinforcement learning.
The paper's approach has been extended by subsequent research, such as fast sampling methods, latent diffusion models, and text-to-image systems. Its success contributed to massive commercial interest, with tools that are now integrated into hardware and software products. The bridge between physics (non-equilibrium thermodynamics, Brownian motion) and modern AI is a testament to academic research.