Jack Rae is a computer scientist specializing in Machine learning and large language models. He is best known for his work at Google DeepMind, where he was a key contributor to the Chinchilla scaling laws, a landmark study that reshaped how researchers approach model and data scaling in deep learning. His research has influenced the design of subsequent generative AI systems and the broader field of artificial intelligence.
Rae's career spans both academic research and industrial application. He has been affiliated with leading AI organizations, including DeepMind and, as of the mid-2020s, Anthropic, where he has continued to work on advancing the capabilities and safety of large language models. His contributions are frequently cited in discussions of efficient training and the trade-offs between model size and dataset size.
Chinchilla Scaling Laws
In 2022, Rae co-authored the paper "Training Compute-Optimal Large Language Models," which introduced the Chinchilla model. The study demonstrated that for a given compute budget, the optimal model size and number of training tokens scale roughly equally, meaning that many existing large models were significantly over-parameterized and under-trained. This finding led to a shift in the field, with subsequent models like Chinchilla being trained on more data with fewer parameters to achieve better performance per unit of compute. The work has become a foundational reference for training transformer-based models, influencing both academic research and industry practices at companies like OpenAI and Anthropic.
Research Contributions
Beyond Chinchilla, Rae has worked on various aspects of neural networks and sequence modeling. His early research included work on sparse attention mechanisms and memory-augmented architectures, which aim to improve the efficiency and long-range dependency handling of transformers. He has also explored topics such as positional encodings and multi-head attention variants, contributing to a deeper understanding of how these components affect model performance. His publications often emphasize empirical analysis and practical improvements over purely theoretical advances.
Career and Affiliations
Rae completed his PhD in computer science at the University of Toronto, where he worked under the supervision of Aaron Courville and others. His doctoral research focused on deep learning for sequential data, laying the groundwork for his later work on language models. After graduating, he joined DeepMind in London, where he became a senior research scientist. At DeepMind, he led or contributed to several high-impact projects, including the Gopher and Chinchilla models. In 2024, he moved to Anthropic, where he has been involved in developing and aligning large language models, with a focus on safety and reliability.
Impact on AI Development
The Chinchilla scaling laws have had a profound impact on the generative AI industry. By highlighting the importance of data efficiency, they encouraged organizations to invest more in data collection and curation rather than simply increasing model size. This has influenced the training strategies of major AI labs, including Google DeepMind, OpenAI, and Anthropic, as well as cloud providers offering AI infrastructure like Amazon Web Services and Google Cloud. The principles from the paper are now standard considerations in designing large language models, affecting everything from learning rate schedules to model pruning techniques.
Selected Works and Recognition
Rae is a co-author of several influential papers, including "Scaling Language Models: Methods, Analysis & Insights from Training Gopher" and "Training Compute-Optimal Large Language Models." His work has been presented at major conferences such as NeurIPS and ICML, and he has been invited to speak at academic and industry events. While he has not received widely publicized awards, his research is highly cited and has been recognized as a key contribution to the field of artificial intelligence. He is also known for his clear communication of complex ideas, often explaining scaling laws and model training in accessible terms for both technical and general audiences.
Future Directions
As of the mid-2020s, Rae continues to work on improving the efficiency and safety of large language models. His current interests include developing methods for better data selection, reducing the environmental impact of training, and ensuring that AI systems align with human values. His move to Anthropic signals a focus on these latter aspects, as the organization is dedicated to AI safety research. Given the rapid evolution of the field, Rae's ongoing contributions are likely to remain influential in shaping how future models are trained and deployed.