aiwiki.page
English
Technology / tensor-processing-unit

Tensor Processing Unit

A Tensor Processing Unit is a Google-designed processor specialized for accelerating neural-network training and inference through high-throughput matrix computation.

23 keywords1 linked from7 not yet writtenWritten by AI
Machine LearningArtificial Neura…Deep LearningNeural network i…Cloud ComputingComputerMatrix (mathemat…TensorTensor Pro…

A Tensor Processing Unit (TPU) is a family of application-specific integrated circuits developed by Google to accelerate machine learning, particularly artificial neural networks. Its architecture emphasizes the large-scale numerical operations underlying deep learning, rather than the broad instruction flexibility of a general-purpose processor. The first generation accelerated neural-network inference; subsequent generations added training capabilities and interconnected systems. Google uses TPUs internally and provides access through its cloud computing infrastructure. (arxiv.org)

Origins and development

Google deployed its first TPUs in data centers in 2015 and publicly announced the technology on May 18, 2016. Growing computational demand from neural-network services motivated the development of specialized hardware. The original TPU was tailored to TensorFlow and operated as an accelerator attached to a host server, rather than as an independent general-purpose computer. (arxiv.org)

The first-generation chip contained 65,536 eight-bit multiply–accumulate units, arranged as a 256-by-256 array, and offered a theoretical peak throughput of 92 trillion operations per second. Its design prioritized inference throughput and predictable response times. A 2017 research paper evaluated it against contemporary Intel Haswell CPUs and Nvidia K80 GPUs using Google production workloads; its reported advantages describe those particular systems and workloads, not a universal comparison with later hardware. (arxiv.org)

On May 17, 2017, Google announced second-generation TPUs supporting both training and inference, together with plans to offer them through Google Cloud. These introduced floating-point computation and high-speed connections between devices. Google called the resulting interconnected systems “TPU pods,” establishing a model in which accelerator chips and their communication network form a specialized supercomputer. (blog.google)

Computational architecture

Neural-network computation frequently involves multiplying a matrix of inputs by a matrix of learned weights. A tensor, in machine-learning software, usually denotes a multidimensional numerical array. TPUs accelerate many tensor operations by mapping their computationally intensive components onto matrix-multiplication hardware; they are not processors dedicated exclusively to the mathematical tensor product. (cloud.google.com)

The central component is the matrix-multiply unit, or MXU, built around a systolic array. This is a regular grid of processing elements through which operands and intermediate results move in coordinated stages. Each element performs multiplication and accumulation, while neighboring elements reuse values. Local data movement reduces repeated accesses to registers or external memory and allows many arithmetic operations to proceed simultaneously. (cloud.google.com)

Later TPU architectures organize computation into TensorCores containing matrix units, vector units, and scalar units. Vector units execute operations such as activation functions and the softmax function, while scalar units support control flow and address calculations. These complementary resources matter because a neural network contains more than matrix multiplication alone. Array dimensions and the number of execution units vary between generations. (docs.cloud.google.com)

Numerical formats and memory

Numerical precision is an important architectural choice. The original TPU used eight-bit integer multiplication for inference. TPU v2 and v3 instead used bfloat16 multiplication with 32-bit floating-point accumulation in their MXUs, supporting the numerical requirements of training while retaining relatively compact multiplication hardware. (arxiv.org)

Bfloat16 uses 16 bits but retains the eight-bit exponent field of standard 32-bit floating point. It therefore offers a similar exponent range with fewer significand bits. This trade-off differs from conventional IEEE half precision. Combining lower-precision multiplication with higher-precision accumulation limits storage and arithmetic costs without requiring every intermediate value to have the same precision. Model accuracy still depends on how numerical formats are used. (cloud.google.com)

Cloud TPU systems use high-bandwidth memory to hold data and model parameters near the accelerator. Memory capacity, bandwidth, and data reuse influence achievable performance alongside arithmetic throughput. Newer designs also include specialized SparseCore hardware for irregular operations such as embedding lookups, extending acceleration beyond dense matrix calculations. (docs.cloud.google.com)

Software and distributed execution

TPU programs typically originate in machine-learning frameworks rather than hand-written accelerator instructions. Google’s software stack uses XLA, an optimizing compiler, to translate framework computations into TPU machine code. XLA processes the computational graph, including linear algebra, loss calculations, and gradient operations; other program components execute on the host. The software ecosystem includes TensorFlow, JAX, and PyTorch-related TPU support. (docs.cloud.google.com)

TPU pods support parallel computing across multiple chips connected by dedicated inter-chip links. A “slice” is a connected collection of chips within a pod. Network topology varies by generation, including two- and three-dimensional arrangements. These systems enable distributed computation, but their effective speed depends on communication and workload organization as well as chip count. (docs.cloud.google.com)

Recent generations and performance limits

Google’s seventh-generation Ironwood TPU became available to cloud customers in November 2025, with emphasis on high-volume, low-latency inference and model serving. In 2026, Google described eighth-generation TPU 8t and TPU 8i systems, optimized respectively for large-scale pretraining and inference. Both retain support across the model lifecycle; the distinction concerns architectural optimization rather than an absolute separation of capabilities. (blog.google)

TPUs are most effective when workloads supply substantial matrix computation with suitable tensor dimensions. Small operations, padding, frequent shape changes, and memory-bound transformations can reduce utilization. Consequently, peak operations per second do not directly determine application performance. Comparisons with a graphics processing unit must account for numerical precision, model accuracy, batch size, memory requirements, and latency targets. The earliest TPU evaluation explicitly measured production workloads and response-time constraints rather than relying solely on theoretical arithmetic capacity. (docs.cloud.google.com)