The US AI Safety Institute (AISI), established in 2023 within the National Institute of Standards and Technology (NIST), conducts research aimed at ensuring the safe development and deployment of advanced Artificial intelligence systems. Its research agenda spans technical evaluations, red-teaming, and the development of measurement standards for AI capabilities and risks. The institute collaborates with academic institutions, industry partners, and international bodies to address challenges posed by Generative AI and other emerging technologies.
AISI's research is grounded in the belief that rigorous, empirical testing is essential to understand AI systems' behavior, limitations, and potential harms. The institute's work informs policy recommendations and best practices for AI developers, with a focus on frontier models that exhibit advanced capabilities in areas such as Large language models and Deep learning.
Technical Evaluations and Red-Teaming
A central component of AISI's research involves designing and conducting technical evaluations of AI models. These evaluations test for dangerous capabilities, including cyber offense, biological threat creation, and autonomous replication. Researchers employ red-team methodologies, where teams attempt to elicit harmful outputs or bypass safety measures. In 2024, AISI published a framework for evaluating frontier AI models, which includes standardized benchmarks and adversarial testing protocols. The institute also develops tools for automated evaluation, enabling scalable assessments of model behavior across diverse scenarios.
Safety Standards and Guidelines
AISI contributes to the development of safety standards and guidelines for AI systems. This includes defining metrics for robustness, fairness, and transparency, as well as establishing best practices for model documentation and deployment. The institute works with standards organizations to create internationally recognized norms, such as the NIST AI Risk Management Framework, which provides a structured approach to managing AI risks. Research in this area also explores the effectiveness of techniques like Model Pruning and Data Augmentation in improving model safety and reliability.
Collaboration and Partnerships
AISI collaborates with leading AI research institutions, including MIT CSAIL, Stanford AI Lab, and BAIR (Berkeley AI Research), as well as industry players like OpenAI, Anthropic, and Google DeepMind. These partnerships facilitate access to cutting-edge models and expertise, enabling AISI to conduct pre-deployment testing of frontier systems. The institute also engages with international counterparts, such as the UK AI Safety Institute, to harmonize evaluation approaches and share findings. In 2024, AISI announced a partnership with Carnegie Mellon University to develop new methods for verifying AI system safety.
Policy and Public Engagement
AISI's research directly informs policy decisions and public discourse on AI safety. The institute publishes reports and briefings that summarize findings and make recommendations for regulators and lawmakers. It also hosts workshops and conferences to foster dialogue among stakeholders, including civil society and academia. AISI's work has influenced executive orders and legislative proposals related to AI oversight, emphasizing the need for evidence-based regulation. The institute maintains a public repository of evaluation results and safety research, promoting transparency and accountability.
Future Directions
Looking ahead, AISI plans to expand its research into areas such as interpretability, alignment, and the societal impacts of AI. The institute is exploring the use of reinforcement-learning-from-human-feedback (RLHF) and other techniques to improve model alignment with human values. Additionally, AISI is investigating the potential risks of Artificial general intelligence and long-term safety concerns. As AI technology evolves, AISI aims to remain at the forefront of safety research, adapting its methods to address new challenges and ensure that AI benefits society while minimizing harms.