Responsible scaling policies (RSPs) are voluntary frameworks published by AI developers that tie increases in a model's deployed capability to corresponding increases in the safety, security, and evaluation measures applied to it, with the explicit goal of preventing catastrophic harm as models approach dangerous capability thresholds. Anthropic published the first such policy in September 2023, defining a series of "AI Safety Levels" (ASL) loosely modeled on biosafety containment standards, and the approach was subsequently adopted in similar form, under different names, by other labs.
The core idea is to make safety commitments concrete and auditable rather than purely aspirational: a lab commits in advance to specific capability evaluations, and states that if a model crosses a defined threshold, for example meaningfully uplifting a novice attempting to create biological weapons, it will not deploy or continue training the model until matching safeguards are in place.
Origins and comparable frameworks
Anthropic's RSP was influenced by earlier AI safety research on measuring and forecasting dangerous capabilities, and by internal debate over how to operationalize alignment concerns into concrete deployment decisions. OpenAI published a comparable "Preparedness Framework" in December 2023, tracking risk categories such as cybersecurity and biological threats through a scorecard model. Google DeepMind followed with a "Frontier Safety Framework" in 2024, and several other developers of frontier models made similar voluntary commitments around the time of the 2023 AI Safety Summit at Bletchley Park.
Content and mechanics
Typical elements include capability thresholds tied to specific evaluations, required safety and security measures at each level, such as red teaming and restricted model-weight access, and commitments to pause scaling if measures cannot be readied in time. Anthropic CEO Dario Amodei has framed the policy as a bet that racing to build powerful AI safely is preferable to not racing at all, given that other actors will continue development regardless. Policies are self-published and self-enforced rather than externally audited or legally binding, a structural feature shared across the industry as of the mid-2020s.
Criticism
Critics, including some AI governance researchers, argue that voluntary industry self-regulation lacks enforcement teeth, that thresholds can be redefined by the same companies obligated to meet them, and that competitive pressure creates an incentive to interpret ambiguous evaluations generously. Supporters counter that RSPs at least establish a public paper trail and an industry norm that binding regulation, such as the EU AI Act, can later reference or formalize, and that they represent one of the few concrete artifacts connecting abstract existential risk concerns to day-to-day deployment decisions.