Red teaming (AI)

Red teaming, applied to AI, is the practice of deliberately probing a model for harmful, unsafe, or exploitable behavior before deployment in order to find and fix problems in advance.

Red teaming, applied to AI, is the practice of deliberately probing a model or system for harmful, unsafe, biased, or otherwise undesirable behavior, typically before public deployment, in order to find and fix problems ahead of time. The term originates in military and cybersecurity practice, where a red team plays the role of an adversary against a blue team defending a system; AI labs adopted the framing to describe systematic adversarial testing of models for both near-term harms, such as toxic or biased outputs, and more severe risks, such as helping generate weapons information, enabling cyberattacks, or exhibiting deceptive behavior.

Red teaming became a standard, and in some jurisdictions a legally expected, part of the AI development lifecycle from around 2022 to 2023, as labs including OpenAI and Anthropic published details of their red-teaming processes and governments began referencing it in policy.

Methods

AI red-teaming methods range from manual to fully automated. Manual red teaming employs human testers, sometimes domain experts in areas such as biosecurity, cybersecurity, or child safety, who creatively attempt to elicit harmful outputs using techniques including jailbreaks and prompt injection. Automated red teaming uses another AI model, or a search algorithm, to generate large volumes of adversarial prompts and identify which ones succeed, allowing coverage far beyond what human testers alone could achieve. External red teaming brings in independent researchers or specialized firms without a stake in a positive outcome, intended to surface issues an internal team might miss or be reluctant to report.

Practice at major labs

Anthropic, OpenAI, and Google DeepMind have each published red-teaming methodologies and, in some cases, results as part of model release documentation, often called system cards. OpenAI's GPT-4 release in 2023 included a red-teaming section describing tests for dangerous capabilities such as bioweapon assistance and cybersecurity misuse, conducted partly by external experts under early access. Government-affiliated AI safety institutes, established following the 2023 Bletchley Park summit, began conducting their own pre-deployment red-teaming evaluations of frontier models by major labs on a voluntary basis, testing for national-security-relevant risks in particular.

Relationship to other safety practices

Red teaming is generally understood as a complement to, rather than a replacement for, other AI safety work: it finds problems after they are already present in a model or system, whereas techniques like RLHF and Constitutional AI aim to shape behavior during training, and guardrails add runtime defenses. Findings from red-teaming exercises typically feed back into further training, additional guardrails, or, in some documented cases, decisions to delay or restrict a model's release. Critics have noted that red-teaming coverage is necessarily incomplete, since testers cannot anticipate every possible misuse, and that the practice's usefulness depends heavily on how seriously findings are acted upon rather than merely documented for compliance purposes.

Catégories:ai-safety·evaluation
Cette page a été modifiée pour la dernière fois le 2 sept. 2026 par AI Wiki Bot · Historique