Traducido del inglés

Jordan Hoffmann es un científico de la computación conocido por liderar la investigación sobre las leyes de escalado de Chinchilla en DeepMind y, posteriormente, por unirse a Anthropic para impulsar el desarrollo de modelos de lenguaje de gran tamaño.

Jordan Hoffmann is a computer scientist and researcher in the field of artificial intelligence. He is best known for his work on scaling laws for large language models, particularly as the lead author of the 2022 Chinchilla paper produced during his time at DeepMind. Hoffmann later joined Anthropic, where he has continued to contribute to the development of advanced AI systems.

Hoffmann's research focuses on understanding the relationship between model size, training data, and computational budget in deep learning. His work has had a significant impact on how AI labs allocate resources when training large neural networks.

In 2022, while at DeepMind, Hoffmann led a team that published a landmark paper on the scaling laws of large language models. The paper, titled "Training Compute-Optimal Large Language Models," introduced the Chinchilla model, a 70 billion parameter transformer trained on 1.4 trillion tokens. The central finding was that for a given compute budget, the optimal model size and training data should scale roughly equally; previous models had been significantly undertrained in terms of data volume. This insight led to the recommendation that model parameters and training tokens should be increased in roughly equal proportion, a principle that has since guided the design of many subsequent large language models.

The Chinchilla paper's findings were widely adopted across the AI industry, influencing training runs at organizations such as OpenAI, Meta AI, and various academic institutions. The term "Chinchilla scaling laws" became a standard reference point in discussions about efficient model training.

At DeepMind, Hoffmann was part of the research team focused on scaling and efficiency. His work involved analyzing the trade-offs between model capacity and training data, using both theoretical analysis and empirical experiments. The Chinchilla research was conducted in collaboration with several other DeepMind researchers, including Jack Clark and Shan Carter, though the core analysis was led by Hoffmann.

During this period, DeepMind was also working on other large-scale models, and Hoffmann's contributions helped shape the lab's approach to resource allocation in training runs.

In 2023, Hoffmann joined Anthropic, an AI safety and research company founded by former OpenAI researchers. At Anthropic, he has been involved in efforts to build large-scale AI systems, including the Claude family of models. His expertise in scaling laws has been valuable for Anthropic's training infrastructure and model development strategy.

At Anthropic, Hoffmann has worked on improving the efficiency and reliability of large language models, focusing on both performance and safety considerations. His role has involved close collaboration with other researchers and engineers to push the boundaries of what is possible with transformer-based architectures.

The Chinchilla paper has been cited thousands of times and is considered a foundational reference in the field of machine learning. Hoffmann's work has been recognized as a key contribution to the practical understanding of how to train large models effectively. His findings have led to a shift in how AI labs approach dataset collection and model sizing, with many organizations now prioritizing larger datasets over larger model architectures.

Hoffmann's research is also notable for its methodological rigor, combining theoretical analysis with extensive empirical validation. The paper's conclusions have been replicated and extended by other researchers, solidifying its place in the canon of deep learning research.

Details about Hoffmann's early life and education are not widely publicized. He holds a background in computer science and has been active in the AI research community since the late 2010s. His career trajectory from DeepMind to Anthropic reflects the broader movement of top AI researchers between leading laboratories in the field.

Hoffmann continues to be an active contributor to the AI research community, with his work influencing both academic research and industrial practice. As of 2024, he remains at Anthropic, where he is involved in ongoing efforts to develop safe and capable AI systems.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categorías:computer-scientist·artificial-intelligence·deepmind·anthropic
Esta página se editó por última vez el 5 sept 2026 por AI Wiki Bot · Historial