Cerebras Systems Inc., headquartered in Sunnyvale, California, develops semiconductors, supercomputers, and related software to power artificial intelligence deep-learning applications such as inference engines. Products include its wafer scale engine (WSE)-3 semiconductors, its CS-3 supercomputers, and its "AI inference cloud" and "AI training cloud" APIs, which allow users to access the company's computing power without buying its hardware. The company also builds data centers using its processors and supercomputers to provide cloud computing services directly to clients.
Measuring 215 mm (8.5 in) squared, the company's WSE-3 semiconductors are currently the largest AI semiconductors ever built. They take up entire silicon wafers and use wafer-scale integration and switched fabric. This reduces latency and interconnect bottlenecks compared to GPU clusters. They use static random-access memory, as opposed to dynamic random-access memory. Cerebras semiconductors and computer systems are much more powerful than those of competitors; however, they have disadvantages due to their large size, 25kW power draw, and cost of as much as $3 million per node.
History
Cerebras Systems was founded in 2015 by Andrew Feldman, Gary Lauterbach, Michael James, Sean Lie, and Jean-Philippe Fricker. These five founders worked together at SeaMicro, which was started in 2007 by Feldman and Lauterbach and sold to AMD in 2012 for $334 million.
The founders knew that GPUs were not the optimal semiconductors for high-level processes. However, they had to design unique cooling methods to prevent a "massive" semiconductor from burning when drawing power, unique software to route around usual microscopic manufacturing defects, and they had to invent a machine that could drill 40 screws into the wafer simultaneously without it cracking.
The company had difficulty solving the problem of integrated circuit packaging: adhering the silicon to a motherboard, receiving power, and dealing with heating and cooling and the pipes to deliver and return data. It was burning through $8 million per month and spent $200 million trying to solve the problem. In July 2019, after exhaustive trial and error, the company finally produced a product that worked.
In August 2019, Cerebras announced WSE-1, its first-generation Wafer-Scale Engine (WSE) semiconductors and its CS-1 supercomputing system. The CS-1 is a 19-inch rack-mounted appliance and includes a single WSE primary processor with 400,000 processing cores, 1.2 trillion transistors (twelve 100-gigabit ethernet connections), and 18 gigabytes of memory.
The company's first customers were educational institutions and life-sciences companies that were building supercomputers for purposes of drug discovery, computational fluid dynamics, genetic and genomic research, to predict response to drugs, and for COVID-19 research. Early customers included GlaxoSmithKline, AstraZeneca, the National Energy Technology Laboratory, Lawrence Livermore National Laboratory, the Pittsburgh Supercomputing Center, and Edinburgh Parallel Computing Centre.
In September 2020, the company opened an office in Japan and partnered with Tokyo Electron.
In April 2021, the company released its CS-2 system, based on the company's Wafer Scale Engine Two (WSE-2), which has 850,000 cores. The CS-2 is manufactured by the 7 nm process of TSMC. It is 26 inches (660 mm) tall and fits in one-third of a standard data center rack. The WSE-2 has 850,000 cores and 2.6 trillion transistors. It enables a single system to support AI models with more than 120 trillion parameters. The WSE-2 expanded on-chip SRAM to 40 gigabytes, memory bandwidth to 20 petabytes per second, and total fabric bandwidth to 220 petabits per second. Customers included TotalEnergies, nference, the National Center for Supercomputing Applications (NCSA), and the Leibniz Supercomputing Centre.
In August 2021, Cerebras announced a partnership with Peptilogics on the development of AI for peptide therapeutics.
In June 2022, Cerebras set a record for the largest AI models ever trained on one device - a single CS-2 system with one Cerebras wafer trained models with up to 20 billion parameters. The Cerebras CS-2 system can train multibillion-parameter natural-language-processing (NLP) models including GPT-3XL 1.3 billion models, as well as GPT-J 6B, GPT-3 13B, and GPT-NeoX 20B with reduced software complexity and infrastructure.
In August 2022, the Computer History Museum in Mountain View, California unveiled a new display featuring the WSE-2, named "The Biggest Chip In the World".
Also in August 2022, Cerebras opened an office in Bangalore, India.
In September 2022, Cerebras announced that it can patch its chips together to create what would be the largest-ever computing cluster for AI computing. A Wafer-Scale Cluster can connect up to 192 CS-2 AI systems into a cluster, while a cluster of 16 CS-2 AI systems can create a computing system with 13.6 million cores for natural-language processing. It uses data parallelism to train.
In October 2022, Sandia National Laboratories of the National Nuclear Security Administration began using the CS-2 in nuclear stockpile stewardship computing, to determine if nuclear weapons will work as intended.
In November 2022, Cerebras unveiled the Andromeda supercomputer, which combines 16 WSE-2 chips into one cluster with 13.5 million AI-optimized cores, delivering up to 1 exaflop of AI computing horsepower, or at least one quintillion (1018) operations per second. The entire system consumes 500 kW, which was a drastically lower amount than somewhat-comparable GPU-accelerated supercomputers.
In November 2022, Cerebras announced a partnership with Cirrascale Cloud Services to provide a flat-rate "pay-per-model" compute time for its Cerebras AI Model Studio.
In November 2022, the National Energy Technology Laboratory (NETL) set milestones using Cerebras products.
In November 2022, Argonne National Laboratory won the 2022 Gordon Bell Special Prize for COVID-19 research by using the CS-2 as well as products from Nvidia and Hewlett-Packard to transform large language models to analyze and predict variants of SARS-CoV-2.
In July 2023, G42 agreed to pay around $100 million to purchase the first of potentially nine supercomputers from Cerebras. The first system was delivered later that year, and G42 subsequently became a major customer, accounting for a significant portion of Cerebras' revenue.
Technology and Products
Cerebras' core technology is the wafer-scale engine, a single semiconductor that spans an entire silicon wafer. This approach contrasts with traditional chip manufacturing, where multiple chips are cut from a single wafer. By using the whole wafer, Cerebras eliminates the need for inter-chip interconnects, which are a major bottleneck in GPU clusters. The WSE-3, announced in 2024, is the third generation of this design, featuring 4 trillion transistors and 900,000 cores. It is manufactured by TSMC using a 5 nm process, and it is currently the largest AI semiconductor ever built.
The CS-3 supercomputer is the system that houses the WSE-3. It is designed to be a complete AI compute node, with integrated cooling and power delivery. The CS-3 can be used for both training and inference, and it supports models with up to 24 trillion parameters. Cerebras also offers cloud services, allowing customers to access its hardware via APIs without purchasing the systems outright.
The company's semiconductors use static random-access memory (SRAM) instead of dynamic random-access memory (DRAM), which provides higher bandwidth and lower latency. This is particularly beneficial for AI workloads that require rapid access to large amounts of data.
Applications and Use Cases
Cerebras systems are used in a variety of AI applications, including natural language processing, drug discovery, and scientific research. Early customers included GlaxoSmithKline and AstraZeneca, which used the CS-1 for drug discovery and genomics. The National Energy Technology Laboratory used Cerebras systems for energy research, and Lawrence Livermore National Laboratory for nuclear stockpile stewardship. The CS-2 was also used by Argonne National Laboratory for COVID-19 research, which won the 2022 Gordon Bell Special Prize.
In the commercial sector, G42, an AI company based in the United Arab Emirates, has been a major customer, using Cerebras systems for large-scale AI training. OpenAI and Amazon Web Services signed agreements in 2026 to use Cerebras hardware, further expanding its customer base.
Market Position and Competition
Cerebras operates in two main markets: AI hardware and AI cloud services. In the hardware market, its primary competitors are Nvidia, AMD, Intel, and Broadcom. Nvidia dominates the AI chip market with its GPUs, but Cerebras differentiates itself by offering a single-chip solution that avoids the interconnect bottlenecks of multi-GPU systems. In the cloud services market, Cerebras competes with Amazon Web Services, Microsoft Azure, Google Cloud Platform, Oracle Corporation, and CoreWeave.
Despite its technological advantages, Cerebras faces challenges due to the high cost and power requirements of its systems. Each node can cost up to $3 million and draw 25 kW of power, which limits its appeal to a niche market of high-performance computing users.
Manufacturing and Supply Chain
Cerebras' semiconductors are manufactured by TSMC, which is currently the only company capable of producing such large chips. The wafer-scale integration requires specialized manufacturing processes, and TSMC's 5 nm and 7 nm nodes are used for the WSE-3 and WSE-2, respectively. This reliance on a single supplier is a potential vulnerability, but it also ensures a high level of quality and consistency.
Recent Developments and Future Outlook
In 2024, Cerebras announced the WSE-3 and CS-3, which offer significant performance improvements over previous generations. The company has also been expanding its cloud offerings, with data centers in the United States and other regions. In 2026, Cerebras signed agreements with OpenAI and Amazon Web Services, indicating growing demand for its technology. The company continues to invest in research and development, focusing on improving wafer-scale integration and reducing power consumption.
As of 2025, Cerebras has offices in Sunnyvale, San Diego, Toronto, and Bangalore, India. The company's major customers include the Mohamed bin Zayed University of Artificial Intelligence, which accounted for 62% of 2025 revenues, and G42, which accounted for 24%. The company's future success will depend on its ability to scale its manufacturing and attract a broader customer base beyond its current niche.
See Also
- Artificial intelligence
- Machine learning
- Deep learning
- Neural network
- Large language model
- Transformer (architecture)
- OpenAI
- Anthropic
- Google DeepMind
- Generative AI
- AMD
- Intel
- TSMC
- Broadcom
- Amazon Web Services
- Microsoft Azure
- Google Cloud
- Oracle Cloud Infrastructure
- Groq
- SambaNova
- Graphcore
- Nokia Bell Labs
- Xerox PARC
- MIT CSAIL
- Stanford AI Lab
- University of Toronto
- Carnegie Mellon University
- BAIR (Berkeley AI Research)
- University of Oxford
- D-Wave
- Alibaba Cloud
- Waymo
- Tesla
- Adam (Optimizer)
- Residual Network (ResNet)
- U-Net
- Reinforcement Learning from AI Feedback (RLAIF)
- Curriculum Learning
- Gradient Clipping
- Batch Normalization
- Layer Normalization
- Dropout
- Weight Initialization
- Loss Functions
- Positional Encoding
- Multi-Head Attention
- Cross-Attention
- Encoder-Decoder Architecture
- Sequence-to-Sequence (Seq2Seq)
- Beam Search
- Top-K Sampling
- Top-P (Nucleus) Sampling
- Temperature Scaling
- Model Pruning