Top-p sampling, also known as nucleus sampling, is a stochastic decoding strategy used to generate sequences from autoregressive probabilistic models, particularly in natural language generation. It was originally proposed by Ari Holtzman, Yejin Choi, and colleagues in 2019 to address the problem of repetitive and nonsensical text produced by deterministic decoding methods like beam search. The technique has since been applied in fields such as protein engineering and geophysics.
In top-p sampling, a probability threshold p is set, and the next item in a sequence is sampled only from the smallest possible set of high-probability candidates whose cumulative probability exceeds p. This method adapts the size of the candidate pool based on the model's certainty, making it more flexible than top-k sampling, which samples from a fixed number of candidates. Due to its effectiveness, top-p sampling is widely used in many Large language model applications.
Technique
At each step of text generation, a language model calculates a probability distribution over its entire vocabulary for the next token. While simply picking the token with the highest probability (greedy search) or a limited set of high-probability sequences (beam search) is possible, these deterministic methods often produce text that is dull, repetitive, or nonsensical. Top-p sampling introduces randomness to avoid these issues while maintaining quality.
The core idea is to sample from a smaller, more credible set of tokens at each step, called the nucleus. This nucleus contains the most likely next tokens whose combined, or cumulative probability, just exceeds the threshold p. By sampling only from this dynamically-sized group, the model can adapt to different situations. When the model is confident about the next token (e.g., one token has a very high probability), the nucleus will be small. When the model is uncertain (the probabilities are more evenly distributed), the nucleus will be larger, allowing for more diversity.
The process at each step is as follows:
- The model calculates the probabilities for all possible next tokens.
- The tokens are sorted by their probability in descending order.
- The nucleus is formed by selecting tokens from the top of the list until their cumulative probability exceeds the predefined threshold, p.
- The probabilities of tokens within this nucleus are then rescaled so that they sum to 1. All tokens outside the nucleus are discarded (given a probability of 0).
- The final next token is randomly sampled from this new, smaller distribution.
Formally, the nucleus, V^(p) ⊆ V, is defined as the smallest set of tokens satisfying: the sum of P(x | x_1, ..., x_{t-1}) for all x in V^(p) is greater than or equal to p. Here, P(x | x_1, ..., x_{t-1}) represents the probability of a token x given the preceding tokens x_1, ..., x_{t-1}.
Example
Imagine at a certain step, a language model has a vocabulary of five words: [the, a, cat, dog, eats] and produces the following probabilities:
- the: 0.5
- a: 0.2
- cat: 0.1
- dog: 0.1
- eats: 0.1
If we set p = 0.8:
- The tokens are sorted by probability: [the, a, cat, dog, eats].
- The cumulative probability is calculated:
- the: 0.5
- the + a: 0.5 + 0.2 = 0.7
- the + a + cat: 0.7 + 0.1 = 0.8
- The nucleus is the smallest set with cumulative probability ≥ 0.8, which is V^(0.8) = {the, a, cat}.
- The probabilities for this set are rescaled to sum to 1:
- P(the) = 0.5 / 0.8 = 0.625
- P(a) = 0.2 / 0.8 = 0.25
- P(cat) = 0.1 / 0.8 = 0.125
- The next token is then sampled from this new distribution, meaning dog and eats have a 0% chance of being chosen.
Top-k sampling
Top-k sampling is a similar technique where the pool of candidate tokens is restricted to the k most likely tokens. The main advantage of top-p is its adaptability. When the model is very certain about the next token (a peaked distribution), the nucleus V^(p) can be very small. When the model is uncertain (a flat distribution), the nucleus can be much larger, allowing for more diversity. In contrast, top-k always samples from a fixed number of tokens, which may be too restrictive or too broad depending on the context.
Applications
While top-p sampling is most famously used as a decoding strategy for large language models, the technique has also been adapted for use in other scientific domains that involve generating or analyzing sequential data from probabilistic models.
Natural language generation
In its original domain of natural language generation, top-p sampling is valued for its ability to produce more diverse and coherent text compared to deterministic methods. It has been shown to be beneficial in tasks like automatic question generation, where sample diversity is important for creating effective training data for question answering models.
Drug and protein design
Top-p sampling is used in computational biology to generate novel molecular and protein sequences from specialized language models. In de novo drug design, chemical language models trained on molecular structures use nucleus sampling to generate focused libraries of new, valid drug candidates. Similarly, in protein engineering, language models trained on protein sequences employ top-p sampling to propose novel sequences with desired properties, expanding the search space beyond natural variants.
Geophysics
In geophysics, top-p sampling has been applied to generate synthetic seismic data or model subsurface structures. By sampling from probabilistic models of geological sequences, researchers can create diverse plausible scenarios that help in uncertainty quantification and interpretation of seismic surveys.
Relationship to Other Decoding Methods
Top-p sampling is one of several stochastic decoding strategies used with autoregressive models. It complements other techniques such as temperature scaling, which adjusts the sharpness of the probability distribution before sampling, and repetition penalty, which discourages the model from repeating tokens. In practice, top-p is often combined with temperature scaling to fine-tune the diversity and quality of generated text. For instance, a lower temperature makes the distribution more peaked, and top-p then selects a smaller nucleus, resulting in more conservative outputs.
Implementation Considerations
In practice, top-p sampling requires sorting the probability distribution at each generation step, which adds computational overhead compared to greedy decoding. However, for many applications, the benefits in output quality outweigh the cost. Efficient implementations in frameworks like PyTorch and TensorFlow optimize this process by using cumulative sum operations and masking. The threshold p is a hyperparameter that can be tuned; common values range from 0.9 to 0.95 for many language generation tasks, but the optimal value depends on the specific model and application.
Impact and Adoption
Since its introduction in 2019, top-p sampling has become a standard component in the decoding strategies of many large language models, including those developed by OpenAI, Anthropic, and Google DeepMind. It is often the default sampling method in text generation APIs and open-source libraries. The technique's adaptability has made it a key tool in balancing creativity and coherence in generative AI systems, influencing how chatbots, content generators, and other applications produce human-like text.