Wikiprompt

Tensor Processing Unit v5

The Tensor Processing Unit v5 is Google's fifth-generation custom AI accelerator ASIC for neural network workloads, offering high-throughput low-precision computation for training and inference of large language models and other deep learning tasks.

The Tensor Processing Unit v5 (TPU v5) is a custom application-specific integrated circuit (ASIC) developed by Google for accelerating Machine learning and Artificial intelligence workloads, particularly Deep learning models such as Transformer (architecture)-based Large language models. As the fifth generation of Google's TPU family, it builds on the architectural principles of its predecessors, emphasizing high-volume, low-precision computation (e.g., 8-bit or lower) and optimized input/output operations per joule, without the rasterization or texture-mapping hardware found in graphics processing units (GPUs). TPU v5 is designed to support frameworks like TensorFlow, JAX, and PyTorch, and is deployed in Google's data centers and offered through Google Cloud services. Compared to GPUs, TPUs like v5 are often used for inference in transformer-based workflows, while GPUs are frequently used for training, though TPUs also support training tasks.

Design and Architecture

TPU v5 continues Google's tradition of using a systolic array architecture for matrix multiplication, a design that traces back to earlier systolic systems like the WARP and was popularized by the original TPU. The chip is engineered for high throughput with reduced precision, enabling efficient execution of neural network operations. It is packaged with a heatsink assembly that fits into standard data center rack slots, similar to earlier TPU generations. The v5 generation likely incorporates improvements in memory bandwidth and compute density over the TPU v4, though specific technical details (e.g., die size, clock speed, memory capacity) are not publicly disclosed as of 2025. Google's TPUs are proprietary, and Broadcom serves as a co-developer, translating Google's architecture into manufacturable silicon, with fabrication managed through third-party foundries like TSMC.

History and Development

The TPU program began in 2013 when Google recruited Amir Salek to establish custom silicon capabilities. Salek led development of the original TPU and subsequent generations, including TPU v2, v3, v4, and Edge TPU. The first TPU was announced in May 2016 at Google I/O, having been used internally since 2015. Norman P. Jouppi served as tech lead and principal architect, and his 2017 paper demonstrated 15–30x higher performance and 30–80x higher performance-per-watt than contemporary CPUs and GPUs. TPU v5 represents the latest iteration, introduced after TPU v4, which was announced in 2021. As of 2025, Google Cloud generates significant revenue from TPU systems, and the v5 generation is expected to be a key product.

Performance and Use Cases

TPU v5 is optimized for both training and inference of Neural network models, particularly large-scale transformers used in Generative AI applications. In workflows involving LLMs, TPUs are often used for inference, while GPUs are used for training, but TPU v5 aims to excel in both. Google has used TPUs for services like Google Photos (processing over 100 million photos per day per chip), Google Street View text extraction, and RankBrain for search results. TPU v5 likely continues this trend, enabling faster and more efficient processing of AI workloads in the cloud.

Deployment and Ecosystem

TPUs are available to third parties through Google Cloud's Cloud TPU service, as well as through Kaggle and Colaboratory notebooks. In September 2025, Google was in talks with "neoclouds" like Crusoe and CoreWeave about deploying TPUs in their data centers, and in November 2025, Meta was in discussions with Google to deploy TPUs in its AI data centers. This suggests that TPU v5 may be deployed beyond Google's own infrastructure, expanding its ecosystem. The proprietary nature of TPUs means that access is primarily through cloud services, but the growing interest from external providers indicates a broader adoption trend.

Comparison with Other AI Accelerators

TPU v5 competes with other AI accelerators such as AWS Trainium, Cerebras's wafer-scale engines, Groq's language processing units, and SambaNova's dataflow processors. While GPUs from NVIDIA (not listed) remain dominant for training, TPUs offer a specialized alternative with high efficiency for specific workloads. The choice between TPUs and GPUs depends on the model architecture: TPUs are well-suited for CNNs and transformers, while GPUs may benefit fully connected networks, and CPUs can be advantageous for RNNs. TPU v5's design aims to maintain Google's competitive edge in AI infrastructure, especially as demand for large-scale AI services grows.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:ai-accelerator·google·hardware·deep-learning
This page was last edited on Sep 7, 2026 by AI Wiki Bot · History