OctoML is a software platform for optimizing and deploying Machine learning models, founded in 2019 by computer scientist Luis Ceze. The company aimed to simplify the process of taking trained models and running them efficiently across various hardware, including cloud instances and edge devices. In 2024, OctoML was acquired by Nvidia, and its technology was integrated into Nvidia's AI software stack.
The platform was built on Apache TVM, an open-source compiler framework for machine learning that Ceze had co-created. OctoML's core offering used automated optimization techniques such as Model Pruning and quantization to reduce model size and latency, making it easier for organizations to deploy Artificial intelligence applications in production.
Founding and Early Development
OctoML was founded in 2019 in Seattle, Washington, by Luis Ceze, who served as CEO. Ceze was a professor at the Paul G. Allen School of Computer Science & Engineering and had previously co-founded Corensic, a startup acquired by F5 Networks in 2012. The initial team included researchers and engineers from the Apache TVM community.
The company raised significant venture capital funding, including a $15 million Series A round in 2019 and a $85 million Series C round in 2021, led by investors such as Amplify Partners, Madrona Venture Group, and Tiger Global. By 2022, OctoML had raised over $130 million in total funding.
Product Evolution and OctoAI
In June 2023, OctoML launched OctoAI, a generative AI product designed to help developers deploy and scale Large language model applications. The product offered managed endpoints for popular open-source models and automated optimization for GPU instances. In January 2024, the company renamed itself to OctoAI to align with its product focus, as CEO Luis Ceze explained the change was to "eliminate potential confusion between our product and corporate name."
In June 2024, OctoAI introduced OctoStack, a platform for customizing and fine-tuning AI models for specific enterprise use cases. The product allowed customers to bring their own data and deploy optimized models on their preferred cloud infrastructure, including Amazon Web Services, Azure, and Google Cloud.
Acquisition by Nvidia
In September 2024, The Information reported that Nvidia was considering acquiring OctoAI for approximately $165 million. On September 25, 2024, Nvidia completed the acquisition. Following the deal, Luis Ceze became a vice president of AI systems software at Nvidia, while retaining his academic position at the University of Washington. The OctoAI team and technology were integrated into Nvidia's AI software division, focusing on optimizing Deep learning models for Nvidia's hardware platforms.
The acquisition was part of Nvidia's broader strategy to strengthen its software ecosystem around its dominant GPU hardware, particularly for Generative AI workloads. OctoAI's optimization tools complemented Nvidia's existing offerings like TensorRT and Triton Inference Server.
Technology and Impact
OctoML's technology leveraged Apache TVM to automatically tune models for specific hardware targets, using techniques like Batch Normalization and Layer Normalization to improve inference performance. The platform supported a wide range of models, including Transformer (architecture)-based architectures, Residual Network (ResNet) CNNs, and U-Net models for image segmentation.
The company's tools were used by thousands of developers and enterprises to reduce inference costs and latency. OctoML also contributed to the open-source community, releasing tools like OctoAI's model serving library and collaborating with academic institutions.
References
[1] Wikipedia: Luis Ceze
[2] OctoML press releases and funding announcements
[3] The Information, September 2024