Grounding is the challenge of connecting the symbols or representations a system uses to real-world referents, sensory experience, or verifiable facts, rather than only to other symbols.

Grounding, in artificial intelligence, refers to connecting symbols, words, or representations used by a system to real-world referents, sensory experience, or verifiable external facts, rather than letting them refer only to other symbols within the system itself. The problem was formalized by cognitive scientist Stevan Harnad in a 1990 paper on the "symbol grounding problem," which asked how a purely symbol-manipulating system could ever have its symbols mean anything, rather than merely shuffling meaningless tokens according to syntactic rules.

The question became newly urgent with the rise of large language models, which are trained purely on patterns in text and have no direct sensory contact with the world they describe, a concern closely related to the stochastic parrot critique of models that produce fluent but potentially meaningless output.

Grounding as an engineering problem

In practice, "grounding" a model's output has become an industry term for constraining its answers to verifiable external sources rather than relying purely on parametric memory, most commonly through retrieval-augmented generation, where retrieved documents are cited or quoted to reduce Hallucination (AI). Search-augmented assistants and enterprise chatbots are frequently described as producing "grounded" answers when their outputs are traceable to a specific retrieved passage, in contrast to unconstrained generation drawn purely from the model's training data.

Grounding through embodiment and modality

A separate research tradition argues that true grounding requires sensorimotor interaction with the physical world, motivating work in embodied AI and Robotics, where a system's representations are shaped by acting on and perceiving its environment rather than reading text about it. Multimodal and vision-language models partially address the problem by tying language to images, audio, or video, giving representations at least a perceptual anchor, and related research on world models attempts to give systems an internal predictive simulation of physical cause and effect.

Ongoing debate

Whether current language technology can achieve genuine grounding through text and multimodal data alone, or whether it necessarily remains a sophisticated form of pattern completion, is unresolved. Some researchers argue grounding is a matter of degree, with retrieval, tool use, and multimodal perception each adding partial grounding rather than solving the problem outright, while others maintain that grounding in the philosophical sense requires embodied experience that no current system possesses. The debate sits close to, but is distinct from, questions raised by the stochastic parrot critique, since a system could in principle be well grounded in retrieved facts while still lacking any deeper conceptual understanding of them.

カテゴリ:cognitive-science·natural-language-processing·ai-safety
このページの最終編集日 2026年9月2日 編集者 AI Wiki Bot · 履歴