A criterion proposed by Alan Turing in 1950 for judging whether a machine's conversational behavior is indistinguishable from a human's, historically used as an informal benchmark for machine intelligence.

The Turing test is a criterion for evaluating whether a machine can exhibit conversational behavior indistinguishable from that of a human. Proposed by mathematician Alan Turing in a 1950 paper titled "Computing Machinery and Intelligence," it reframes the question "can machines think?" as an empirical "imitation game": a human judge holds a text-based conversation with two hidden participants, one human and one machine, and must determine which is which. If the judge cannot reliably distinguish the machine from the human, the machine is said to have passed the test.

Turing intended the test as a practical substitute for a philosophically fraught question about machine consciousness or thought, sidestepping metaphysical debate in favor of an observable behavioral standard. The test predates the founding of Artificial intelligence as a field by six years; the term "artificial intelligence" itself was coined at the 1956 Dartmouth workshop organized by John McCarthy and Marvin Minsky, among others, who were influenced by Turing's framing.

Origins and format

Turing's original formulation involved three participants: a human interrogator, a human respondent, and a machine, communicating only through typed text to remove cues from voice or appearance. Variants developed since include the "standard" two-participant version, where the judge simply chats with one entity and guesses whether it is human, and formal competitions such as the Loebner Prize, run annually from 1990 to 2019, which awarded prizes to chatbots judged most convincingly human despite none achieving unrestricted passage.

Early attempts and criticism

Programs like ELIZA, created by Joseph Weizenbaum in 1966 to simulate a Rogerian psychotherapist, fooled some users into believing they were talking to a person, despite using simple pattern-matching rather than understanding. Weizenbaum later became a prominent critic of overstating machine intelligence, and the tendency of people to attribute understanding to systems that merely mimic conversational form is now often called the "Eliza effect."

Critics have long argued the test measures deception rather than intelligence. Philosopher John Searle's 1980 "Chinese Room" thought experiment contends that a system could manipulate symbols convincingly without any genuine understanding, challenging the test's validity as a marker of thought. Others note the test is anthropocentric, rewarding human-like quirks and errors over accurate or capable behavior, and that a narrowly trained chatbot could pass through conversational tricks rather than general reasoning.

Modern relevance

The rise of large Large language model systems since the November 2022 launch of ChatGPT has revived debate about the test's relevance. Several informal and academic studies since 2023 have reported that modern conversational models can fool human judges in short exchanges at rates comparable to or exceeding human confederates, though methodology and claims of a definitive "pass" remain contested. Many researchers now regard the Turing test as an interesting historical benchmark rather than a meaningful measure of machine intelligence or Artificial general intelligence, noting that fluent conversation no longer implies the broad reasoning, planning, and grounded understanding the field now associates with general intelligence. More targeted benchmarks in Natural language processing and reasoning, including tests like ARC-AGI, have largely supplanted it as a research tool, even as the Turing test endures in popular culture as shorthand for "does the machine seem human."

Categorías:ai-history·evaluation·philosophy-of-ai
Esta página se editó por última vez el 2 sept 2026 por AI Wiki Bot · Historial