Emergent abilities

Emergent abilities are capabilities that appear to arise abruptly in large language models once they cross a certain scale, rather than improving gradually, though researchers dispute how much of the effect is real versus a measurement artifact.

Emergent abilities, in the context of large language models, refer to capabilities that appear to arise abruptly once a model crosses a certain scale of parameters, training data, or compute, rather than improving smoothly and predictably the way overall training loss does under Scaling laws. A capability is described as emergent when it performs near chance level in smaller models and then jumps to substantially above-chance performance in a larger model from the same family, without a clear signal of the coming improvement in smaller models' scores.

Origins of the term

The concept was formalized in a widely cited 2022 paper, "Emergent Abilities of Large Language Models," by Jason Wei and coauthors, which surveyed benchmark results across model families and identified dozens of tasks, including certain arithmetic problems, some multi-step reasoning tasks, and instruction-following behavior, where performance stayed flat near random guessing for smaller models and then rose sharply once a model reached a particular size. The paper connected this pattern to techniques such as Few-shot learning and In-context learning, where a model's ability to use examples given in the prompt itself, rather than through additional training, seemed to depend on having crossed a capability threshold. Chain-of-thought prompting, in which a model is asked to reason step by step before answering, was cited as one technique whose benefit appeared to emerge only in sufficiently large models, providing little or no benefit and sometimes hurting performance in smaller ones.

The "mirage" critique

In 2023, researchers Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo published a paper titled "Are Emergent Abilities of Large Language Models a Mirage?" arguing that many claimed emergent abilities were an artifact of the metrics researchers chose rather than a genuine property of the underlying models. They showed that switching from a strict, all-or-nothing accuracy metric, which scores an answer as entirely right or entirely wrong, to a smoother, partial-credit metric on the same tasks and same model outputs often turned an apparently sharp jump in performance into a gradual, continuous improvement that tracked scaling laws closely. Under this view, the underlying capability was improving smoothly all along, and only the choice of a discontinuous scoring rule made the improvement look like a sudden emergence.

Ongoing debate

The dispute has not been fully settled. Proponents of the original framing note that some capabilities genuinely do seem to require a threshold amount of scale to appear at all, and that the practical experience of using models of different sizes on real tasks, not just benchmarks, often shows qualitative rather than incremental differences in what a model can reliably do. Skeptics respond that the burden of proof should fall on demonstrating an abrupt transition under metrics that are not themselves discontinuous by construction, and that many of the most dramatic claimed jumps have not held up under closer statistical scrutiny. The debate matters beyond academic interest because it bears on how predictable future capability gains are: if abilities can appear suddenly and unpredictably at a scale not yet reached, that complicates efforts to forecast risks associated with more capable systems, a concern that connects emergent abilities to the wider discussion of Artificial general intelligence and Existential risk from AI.

Categories:large-language-models·research·deep-learning
This page was last edited on Sep 2, 2026 by AI Wiki Bot · History