# Scale (Scale AI)

Scale AI, Inc. is an American AI infrastructure and software company based in San Francisco, California, offering data annotation, RLHF services, LLM evaluation, and enterprise AI deployment tools. Founded in 2016, it became a major AI data and evaluation provider for commercial and government clients.

Scale AI, Inc. is an American artificial intelligence infrastructure and software company based in San Francisco, California. Originally focused on data annotation, the company also offers reinforcement learning from human feedback (RLHF) services, large language model (LLM) evaluation, and enterprise software suites to build and deploy AI applications. Its research arm, the Safety, Evaluation and Alignment Lab, focuses on evaluating and aligning LLMs, and it co-created the benchmark Humanity's Last Exam.

Scale AI outsources data labeling through its subsidiaries, Remotasks, which focuses on computer vision and autonomous vehicles, and Outlier, which focuses on annotating data for LLMs. The company operates an LLM Red Team that conducts human adversarial testing to identify vulnerabilities, biases, and safety risks in AI models, working with organizations such as [OpenAI](https://www.wikiprompt.org/wiki/openai) and [Google DeepMind](https://www.wikiprompt.org/wiki/google-deepmind) and national AI Safety Institutes. Commercial customers have included Google, Microsoft, Meta, General Motors, OpenAI, and Time, while the company also works with governments, including the United States on military-related projects and Qatar on social program efficiency.

## History

### Early years (2016–2019)
Scale was founded in 2016 by Alexandr Wang and Lucy Guo through Y Combinator, after the pair had worked together at Quora. Initial investors included Dragoneer Investment Group, Tiger Global Management, and Index Ventures. Guo was fired in 2018. In August 2019, after Peter Thiel’s Founders Fund made a $100 million investment, Scale's valuation exceeded $1 billion, granting it unicorn status.

### Growth (2019–2025)
Scale contracted with the United States Department of Defense in 2020. In May 2021, Michael Kratsios, former Chief Technology Officer of the United States under the Trump administration, joined as managing director and head of strategy. By July 2021, Scale reached a valuation of $7 billion after a financing led by Greenoaks, Dragoneer, and Tiger Global, driven by increased demand for data labeling across industries.

In January 2022, Scale won a $250 million contract to give American federal agencies access to its suite of tools. In February 2022, it developed the Automated Damage Identification Service in response to the Russian invasion of Ukraine, analyzing satellite imagery to measure building damage and geotagging reports for humanitarian groups. In November 2022, Scale was recognized by Time on its Best Inventions of 2022 list and opened an office in St. Louis.

In January 2023, Scale laid off 20% of its workforce. In May 2023, it signed a deal with the US Army’s XVIII Airborne Corps, becoming the first AI company to deploy its LLM, known as Donovan, on a classified network. In August 2023, Scale became OpenAI’s preferred partner to fine-tune GPT-3.5, and its services were used to create ChatGPT. That same month, Scale’s evaluation platform was used at DEF CON for the first generative AI red team event, testing models from various companies. In December 2023, Scale contributed to Meta’s Purple Llama initiative, a security framework for open generative AI models.

In February 2024, Scale was selected by the Department of Defense to test and evaluate its LLMs for military purposes under a one-year contract. In March 2024, Scale reached a valuation of nearly $13 billion after Accel led another funding round. In May 2024, it raised an additional $1 billion with new investors including Amazon and Meta, bringing its valuation to $14 billion. In August 2024, Scale signed an agreement with the US AI Safety Institute, controlled by the Department of Commerce’s National Institute of Standards and Technology, to collaborate on research, testing, and evaluation.

In December 2024, Scale was sued by a former employee alleging wage theft and worker misclassification; a second employee filed a similar suit the following month. In January 2025, several contractors sued Scale alleging psychological harm from exposure to disturbing content. That month, The Conversation reported that Scale and Meta had previously teamed up to create and sell Defense Llama, an LLM product for military-style defense purposes. Scale also took out a full-page ad in The Washington Post appealing to President Donald Trump to "win the AI war". Later in January, Scale and the Center for AI Safety released Humanity's Last Exam, a benchmark for AI systems, and Scale assisted in developing benchmarks EnigmaEval, MultiChallenge, and MASK.

In February 2025, Scale agreed to a five-year partnership with the Qatari government to improve government services via AI tools and training, including predictive analytics, automation, and advanced data analytics, signed at the Web Qatar 2025 Summit. Also in February, Scale became a third-party evaluator of AI models for the US AI Safety Institute. In March 2025, Scale reached a multimillion-dollar deal with the Department of Defense to develop the Thunderforge project, a major step in US military automation, aiming to use AI to plan and help execute movements of ships, planes, and other assets. The contract, awarded by the Defense Innovation Unit to Scale, Anduril Industries, and Microsoft, was intended for initial use with USINDOPACOM and EUCOM. In April 2025, Scale released Scale Evaluation, a platform for testing LLMs against benchmarks to pinpoint weaknesses and flag where additional training data would improve models.

### Additional investment from Meta (2025)
On June 10, 2025, Meta Platforms agreed to purchase a 49% non-voting stake in Scale AI for $14.8 billion. The company remained a standalone, independent entity from Meta. Former CEO Alexandr Wang took a top position inside Meta as part of the deal and was replaced by Jason Droege, the company's chief strategy officer and former Uber executive.

### Scale Labs
In March 2026, the company launched Scale Lab, an initiative aimed at advancing AI research and development, though specific details of its operations were not widely publicized as of that time.

## Services and Products
Scale AI provides a suite of tools for data annotation, including image, video, and text labeling, which are essential for training [machine learning](https://www.wikiprompt.org/wiki/machine-learning) models. Its RLHF services help align LLMs with human preferences, a critical step in developing systems like [large language models](https://www.wikiprompt.org/wiki/large-language-model). The company's evaluation platforms, such as Scale Evaluation, allow organizations to test model performance against benchmarks, identify weaknesses, and improve training data. For enterprise clients, Scale offers software suites to build and deploy AI applications, including tools for model monitoring and safety testing.

The LLM Red Team conducts adversarial testing to uncover vulnerabilities, biases, and safety risks, often simulating complex threats like cybersecurity attacks, jailbreaks, and agentic AI behaviors. This work supports both commercial clients and government agencies, including national AI Safety Institutes.

## Government and Military Work
Scale AI has a significant portfolio of government contracts. Its work with the US Department of Defense includes the Thunderforge project, which aims to automate military planning and execution, and earlier contracts to test and evaluate LLMs for military purposes. The company has also partnered with the Qatari government to enhance social programs through AI, and with the US AI Safety Institute for model evaluation. These collaborations have drawn attention due to ethical concerns about AI in military contexts, but Scale has positioned itself as a leader in safe and aligned AI deployment.

## Research and Benchmarks
Scale's Safety, Evaluation and Alignment Lab focuses on evaluating and aligning LLMs, contributing to the development of benchmarks like Humanity's Last Exam, which tests AI systems across a wide range of knowledge and reasoning tasks. The company has also assisted in creating EnigmaEval, MultiChallenge, and MASK, which target specific capabilities such as reasoning, multi-step problem solving, and safety. These benchmarks are used by researchers and developers to assess model performance and guide improvements.

## Legal and Ethical Issues
Scale has faced legal challenges related to labor practices. In December 2024, a former employee sued the company for wage theft and worker misclassification, followed by a similar suit in January 2025. Additionally, contractors sued Scale in January 2025, alleging psychological harm from exposure to disturbing content during data labeling tasks. These cases highlight ongoing concerns about the working conditions of data annotators, who often handle sensitive or traumatic material. Scale has not publicly settled these cases as of early 2026.

## Leadership and Structure
Alexandr Wang co-founded Scale and served as CEO until June 2025, when he left to join Meta as part of the investment deal. Jason Droege, who had been chief strategy officer and previously worked at Uber, succeeded him as CEO. The company is headquartered in San Francisco, California, and operates subsidiaries Remotasks and Outlier for data labeling. Scale's board and leadership have included figures like Michael Kratsios, who joined in 2021 to lead strategy.

---
Source: https://www.wikiprompt.org/wiki/scale
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-10-07T16:35:21.409587+00:00
