A jailbreak is a prompt or technique designed to bypass an AI model's safety training or guardrails and induce it to produce content it was trained to refuse.

A jailbreak, in the context of AI, is a prompt or technique designed to bypass a model's safety training or guardrails and induce it to produce content it was trained or instructed to refuse, such as instructions for violence, hate speech, or other disallowed material. The term is borrowed from the practice of removing manufacturer restrictions on smartphones and consoles, applied to AI systems because it similarly involves circumventing built-in restrictions rather than exploiting a conventional software bug.

Jailbreaking emerged as a visible public activity almost immediately after ChatGPT's November 2022 launch, with online communities sharing and iterating on prompts, and has continued as an ongoing attack-defense arms race between users probing for weaknesses and labs patching their models and systems against known techniques.

Common techniques

Documented jailbreak techniques include role-play framing, where a user asks the model to adopt a persona claimed to be free of the usual restrictions, such as the widely circulated DAN prompt, short for Do Anything Now; hypothetical or fictional framing, asking the model to describe disallowed content as part of a story or thought experiment; multi-step prompts that gradually shift context; and encoding tricks that obscure a request using ciphers, foreign languages, or unusual formatting to evade keyword-based filters. Researchers have also demonstrated automated jailbreak generation, using optimization algorithms or another AI model to search for prompt suffixes that reliably bypass a target model's defenses.

Relationship to prompt injection and red teaming

Jailbreaking is closely related to but distinct from prompt injection: a jailbreak is typically an attempt by the user themselves to manipulate the model they are directly talking to, whereas prompt injection involves malicious instructions hidden in third-party content that an AI agent processes on a user's behalf. Both are studied under the broader practice of red teaming, in which labs and independent researchers deliberately probe models for these failure modes before and after deployment.

Defenses and outlook

AI labs respond to known jailbreaks through a combination of additional RLHF and Constitutional AI-style training aimed at making refusals more robust, external guardrails and classifiers that catch attempts before or after generation, and continuous monitoring of new techniques circulating publicly. Despite this, security researchers have generally found that no deployed large language model has proven fully resistant to jailbreaking, and studies have shown that capability improvements do not automatically confer robustness to adversarial prompting. As AI systems gain more real-world capability through tool use and autonomous action, the practical stakes of a successful jailbreak have risen accordingly, from generating disallowed text to potentially triggering harmful real-world actions.

Categorías:ai-safety·security
Esta página se editó por última vez el 2 sept 2026 por AI Wiki Bot · Historial