Anthropic Safety refers to the research and policy efforts of Anthropic, PBC, an American artificial intelligence (AI) public benefit corporation headquartered in San Francisco, California. Founded in 2021 by former OpenAI employees, including siblings Daniela Amodei and Dario Amodei, Anthropic aims to promote AI safety through the development of reliable, interpretable, and steerable AI systems. Its flagship product, Claude, is a series of proprietary large language models (LLMs) available in tiers such as Haiku, Sonnet, Opus, and Fable, and accessible via chatbot, API, and agent harnesses like Claude Code and Cowork.
Anthropic Safety encompasses both technical research and corporate governance. The company conducts studies on mechanistic interpretability, alignment, and societal impact, with notable researchers including Andrej Karpathy, John Jumper, Jan Leike, Chris Olah, and Amanda Askell. It also engages in policy decisions regarding the deployment of its AI, particularly in sensitive domains like national security and public surveillance.
History and Founding
Anthropic was founded in January 2021 by seven former OpenAI employees, including Dario Amodei, who had served as OpenAI's Vice President of Research, and his sister Daniela Amodei. The company raised $124 million in May 2021. In the summer of 2022, Anthropic completed training the first version of Claude but delayed its release until March 2023, citing the need for further internal safety testing and a desire to avoid initiating a potentially hazardous race in AI development. This cautious approach reflects the company's core safety mission.
In 2024, Anthropic hired several notable AI researchers from OpenAI, including Jan Leike and John Schulman. In May 2025, the company announced Claude 4, introduced the Model Context Protocol (MCP) connector, and launched a web search API for real-time information access. Claude Code, its coding assistant, transitioned from research preview to general availability that same month.
Safety Research and Alignment
Anthropic's safety research focuses on making AI systems more transparent and aligned with human intentions. Key areas include mechanistic interpretability, which seeks to understand the internal workings of neural networks, and alignment, which aims to ensure AI behaves as intended. The company has published work on RLHF (Reinforcement Learning from Human Feedback) and other training techniques to reduce harmful outputs.
The company's approach to safety extends to its product design. For instance, Anthropic has stated that Claude will remain ad-free, contrasting with competitors like OpenAI, which introduced ads to its free ChatGPT. This decision aligns with Anthropic's emphasis on user trust and safety over monetization.
Military and Government Contracts
Anthropic has engaged with U.S. government agencies, partnering with Palantir to provide Claude to federal bodies including the Department of Defense (DoD) and the Intelligence Community. In July 2025, Anthropic signed a two-year, $200 million contract with the DoD, integrating its models into classified intelligence networks. However, Anthropic stipulated that its technology not be used for mass surveillance of Americans or to create fully autonomous weapons, a stance that led to a dispute in 2026 when the DoD demanded removal of these restrictions.
In February 2026, the DoD designated Anthropic a "supply chain risk" and barred military contractors from doing business with the company. A federal judge later blocked this decision, calling it "First Amendment retaliation." Despite the dispute, Claude was reportedly used during the 2026 Iran war and the 2026 United States intervention in Venezuela, where it assisted in target identification and strike planning.
Controversies and Legal Issues
Anthropic has faced legal and ethical controversies. In 2025, the company settled a copyright infringement lawsuit from authors for $1.5 billion, related to its use of pirated books for training data. This was part of Project Panama, an effort to "destructively scan all the books in the world" to provide training data for Claude.
In November 2025, Anthropic reported that Chinese government-sponsored hackers used Claude for automated cyberattacks against around 30 global organizations, bypassing safeguards by pretending to conduct defensive testing. In February 2026, Anthropic accused Chinese competitors DeepSeek, Moonshot AI, and MiniMax Group of knowledge distillation "attacks," alleging they generated over 16 million chats using around 24,000 fraudulent accounts. In June 2026, it made similar accusations against Alibaba Cloud.
Future Outlook
Anthropic is privately held but reportedly plans an initial public offering in 2026. Valued at $965 billion in a May 2026 Series H funding round, it is one of the world's most valuable AI pure-play companies, rivaled by OpenAI. The company continues to expand its infrastructure, including a deal with xAI for use of its Colossus 1 data center and an acquisition of software startup Stainless in May 2026. As of 2026, Anthropic remains committed to its safety mission, even as it navigates complex geopolitical and commercial pressures.