# AI Safety Institute

The AI Safety Institute is a UK government-backed organization established in November 2023 to evaluate and ensure the safety of advanced artificial intelligence models, later renamed the AI Security Institute in 2025.

The AI Safety Institute is a state-backed organization established to evaluate and ensure the safety of advanced artificial intelligence (AI) models, often referred to as frontier AI models. It was created by the United Kingdom government in November 2023, evolving from the Frontier AI Taskforce, following the AI Safety Summit hosted in the UK. The institute's primary mission is to conduct independent safety evaluations of cutting-edge AI systems, aiming to set international standards for AI safety and governance.

The institute gained prominence during a period of heightened concern about AI risks, particularly after public declarations in 2023 about potential existential threats from advanced AI. The UK government, under Prime Minister Rishi Sunak, emphasized the need for independent oversight, with Sunak stating that AI companies cannot 'mark their own homework'. This led to the establishment of the institute as a dedicated body for pre-deployment testing and safety research.

## Establishment and Early Activities

The AI Safety Institute was officially launched during the AI Safety Summit in November 2023, succeeding the Frontier AI Taskforce. Initially based in London, it announced plans in May 2024 to open an office in San Francisco, where many leading AI companies are headquartered. This move was part of a broader strategy to 'set new, international standards on AI safety', according to UK technology minister Michele Donelan.

In April 2024, the institute concluded an agreement with its US counterpart to collaborate on joint safety tests. This collaboration aimed to harmonize evaluation procedures, as many AI companies had been reluctant to share pre-deployment access to their most advanced models, waiting for common rules to be established. The institute's work focuses on evaluating models from major developers, including those from [openai](https://www.wikiprompt.org/wiki/openai), [anthropic](https://www.wikiprompt.org/wiki/anthropic), and [google-deepmind](https://www.wikiprompt.org/wiki/google-deepmind).

## International Network

The AI Safety Institute became part of a broader international network of similar bodies following the AI Seoul Summit in May 2024. International leaders agreed to form a network comprising institutes from the UK, US, Japan, France, Germany, Italy, Singapore, South Korea, Australia, Canada, and the European Union. This network aims to coordinate safety evaluations and share best practices across jurisdictions.

In July 2025, the network conducted an exercise to explore issues with evaluating AI agents, particularly concerning sensitive information leakage and cybersecurity. Network members also met at NeurIPS 2025 in San Diego to further collaborative efforts. The network's formation reflects growing recognition that AI safety requires international cooperation, given the global nature of AI development and deployment.

## Evaluation and Research Focus

The institute's core work involves conducting independent evaluations of frontier AI models before their public release. This includes assessing capabilities in areas such as [large-language-model](https://www.wikiprompt.org/wiki/large-language-model) performance, potential for misuse, and alignment with human values. The institute employs technical experts who design and run rigorous tests, often using methodologies from [machine-learning](https://www.wikiprompt.org/wiki/machine-learning) research.

A key focus is on identifying and mitigating risks associated with advanced AI systems, including those related to [generative-ai](https://www.wikiprompt.org/wiki/generative-ai) and [deep-learning](https://www.wikiprompt.org/wiki/deep-learning) technologies. The institute also researches safety mechanisms like [rlaif](https://www.wikiprompt.org/wiki/rlaif) and [model-pruning](https://www.wikiprompt.org/wiki/model-pruning) to understand how these techniques affect model behavior. Its findings inform government policy and contribute to the development of safety standards that can be adopted internationally.

## Renaming and Evolution

In 2025, the UK's AI Safety Institute was renamed the 'AI Security Institute', reflecting a broadening of its mandate to include security concerns alongside safety. This change mirrored a similar evolution in the US, where the counterpart became the Center for AI Standards and Innovation (CAISI). The renaming signaled a shift toward addressing not only accidental harms but also deliberate misuse of AI systems.

Despite the name change, the institute continues its core mission of evaluating frontier AI models and promoting safe development practices. It maintains close ties with academic institutions and industry partners, including [mit-csail](https://www.wikiprompt.org/wiki/mit-csail) and [stanford-ai-lab](https://www.wikiprompt.org/wiki/stanford-ai-lab), to stay at the forefront of AI safety research. The institute's work remains crucial as AI capabilities continue to advance rapidly.

## Impact and Legacy

The AI Safety Institute has played a significant role in shaping global AI governance. Its establishment marked a milestone in government involvement in AI oversight, moving beyond voluntary industry self-regulation. The institute's model of independent evaluation has been adopted by other countries, contributing to the creation of similar bodies worldwide.

The institute's emphasis on pre-deployment testing has influenced how AI developers approach safety, encouraging more rigorous internal evaluation processes. Its collaborative efforts with international partners have helped create a more coordinated global response to AI risks. As of 2025, the institute continues to operate as a key player in the evolving landscape of AI safety and security.

---
Source: https://www.wikiprompt.org/wiki/ai-safety-institute
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-09T01:57:11.670736+00:00
