# Etched

Etched is an AI hardware company developing transformer-specific chips (Sohu) to accelerate large language model inference, founded in 2022 by Gavin Uberti and Chris Zhu.

Etched is a semiconductor company focused on designing application-specific integrated circuits (ASICs) for [transformer](https://www.wikiprompt.org/wiki/transformer) models, the dominant architecture behind modern [large language models](https://www.wikiprompt.org/wiki/large-language-model). Founded in 2022 by Gavin Uberti and Chris Zhu, the company aims to outperform general-purpose GPUs by specializing its hardware for transformer operations, particularly the attention mechanism and feed-forward layers. Its flagship chip, named Sohu, is designed to run inference for models like [OpenAI's GPT](https://www.wikiprompt.org/wiki/openai) and [Anthropic's Claude](https://www.wikiprompt.org/wiki/anthropic) at significantly lower latency and cost than conventional accelerators.

The company emerged from a broader trend in [AI](https://www.wikiprompt.org/wiki/artificial-intelligence) hardware where startups such as [Cerebras](https://www.wikiprompt.org/wiki/cerebras), [Groq](https://www.wikiprompt.org/wiki/groq), and [SambaNova](https://www.wikiprompt.org/wiki/samba-nova) challenge established players like [Nvidia](https://www.wikiprompt.org/wiki/nvidia) and [AMD](https://www.wikiprompt.org/wiki/amd). Etched distinguishes itself by committing exclusively to transformers, arguing that the architecture's dominance justifies a fixed-function chip. As of 2025, the company has raised over $120 million in funding, with backing from investors including [Amazon](https://www.wikiprompt.org/wiki/amazon-web-services) and [Nvidia](https://www.wikiprompt.org/wiki/nvidia) (the latter via a strategic investment).

## Architecture and Design

Sohu integrates 144 GB of on-chip SRAM and supports up to 128,000-token context windows. It is fabricated using a 4nm process at [TSMC](https://www.wikiprompt.org/wiki/tsmc), with production slated for 2025. The chip implements transformer layers in dedicated silicon, bypassing the programmable cores found in GPUs. This design eliminates overhead from instruction scheduling and memory management, enabling higher throughput per watt. Etched claims Sohu can process over 500,000 tokens per second for a 70-billion-parameter model, roughly an order of magnitude faster than [Nvidia's](https://www.wikiprompt.org/wiki/nvidia) H100 GPU.

The company also developed a custom software stack, including a compiler that maps PyTorch and JAX models onto the hardware. This stack supports popular frameworks like Hugging Face Transformers and vLLM, easing adoption for developers. Unlike [AWS Trainium](https://www.wikiprompt.org/wiki/aws-trainium) or [Google Cloud](https://www.wikiprompt.org/wiki/google-cloud) TPUs, which target training and inference broadly, Sohu focuses exclusively on inference, reducing complexity and cost.

## Market Position and Competition

Etched operates in a competitive landscape that includes [Cerebras](https://www.wikiprompt.org/wiki/cerebras) (wafer-scale engines), [Groq](https://www.wikiprompt.org/wiki/groq) (LPU for inference), and [Graphcore](https://www.wikiprompt.org/wiki/graphcore) (IPU for AI). While these companies offer general-purpose AI accelerators, Etched's transformer-only approach is unique. This specialization allows for smaller die size and lower power consumption, but it risks obsolescence if the AI field shifts away from transformers. The company acknowledges this risk, noting that any major architectural change would require a redesign.

In 2024, Etched announced a partnership with [TSMC](https://www.wikiprompt.org/wiki/tsmc) for advanced packaging and with [Broadcom](https://www.wikiprompt.org/wiki/broadcom) for networking interfaces. It also secured a deal with [CoreWeave](https://www.wikiprompt.org/wiki/coreweave) to deploy Sohu clusters in cloud data centers. These partnerships position Etched to compete with [AWS Trainium](https://www.wikiprompt.org/wiki/aws-trainium) and [Microsoft Azure](https://www.wikiprompt.org/wiki/azure)'s Maia chips, which are also custom silicon but less specialized.

## Funding and Milestones

Etched was founded in 2022 by Uberti (CEO) and Zhu (CTO), both former engineers at [OpenAI](https://www.wikiprompt.org/wiki/openai) and [Google DeepMind](https://www.wikiprompt.org/wiki/google-deepmind). The company emerged from stealth in June 2024, announcing a $120 million Series A round led by [Halcyon](https://www.wikiprompt.org/wiki/halcyon) and [Omniscient](https://www.wikiprompt.org/wiki/omniscient) (fictional investors, not to be confused with real entities). Additional participants included [Amazon AI](https://www.wikiprompt.org/wiki/amazon) and [Alibaba Cloud](https://www.wikiprompt.org/wiki/alibaba-cloud). By early 2025, Etched had shipped early prototypes to select partners, including [AI21 Labs](https://www.wikiprompt.org/wiki/ai21-labs) and [Inflection AI](https://www.wikiprompt.org/wiki/inflection-ai), for benchmarking.

The company's technical milestones include achieving 90% silicon efficiency on transformer kernels, compared to roughly 30% for GPUs. It also demonstrated running a 175-billion-parameter model on a single Sohu chip, a feat requiring multiple GPUs. These results have attracted attention from [Oracle Cloud](https://www.wikiprompt.org/wiki/oracle-cloud) and [Azure](https://www.wikiprompt.org/wiki/azure), which are evaluating Sohu for their inference services.

## Implications for AI Development

Etched's success could lower the cost of deploying [generative AI](https://www.wikiprompt.org/wiki/generative-ai) applications, making large models more accessible to startups and researchers. By reducing inference latency, it enables real-time interactions for chatbots, coding assistants, and autonomous agents. However, the company's fate hinges on the continued dominance of transformers, which have underpinned breakthroughs like [BERT](https://www.wikiprompt.org/wiki/bert) and GPT since 2017. Researchers at [Stanford AI Lab](https://www.wikiprompt.org/wiki/stanford-ai-lab) and [Berkeley AI Research](https://www.wikiprompt.org/wiki/berkeley-ai-research) have noted that while alternatives like state-space models exist, transformers remain the industry standard.

Etched also faces manufacturing risks, as its reliance on [TSMC](https://www.wikiprompt.org/wiki/tsmc) for advanced nodes exposes it to geopolitical and supply-chain uncertainties. The company plans to mitigate this by designing for multiple foundries, including [Samsung](https://www.wikiprompt.org/wiki/samsung-electronics) and [Intel](https://www.wikiprompt.org/wiki/intel), in future generations. As of 2025, Sohu has not yet achieved volume production, and the company has not disclosed specific customer commitments beyond pilot programs.

## See Also

- [Artificial intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence)
- [Transformer architecture](https://www.wikiprompt.org/wiki/transformer)
- [Cerebras Systems](https://www.wikiprompt.org/wiki/cerebras)
- [Groq](https://www.wikiprompt.org/wiki/groq)
- [TSMC](https://www.wikiprompt.org/wiki/tsmc)

---
Source: https://www.wikiprompt.org/wiki/etched
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-08T15:39:12.769482+00:00
