A graphics processing unit (GPU) is a processor designed to perform many related calculations concurrently, originally for generating images and now also for general-purpose computation. Its architecture emphasizes throughput: the amount of work completed over time. GPUs combine programmable execution units with specialized hardware for particular operations, making them useful for graphics rendering, scientific applications, and machine learning. They typically work alongside a central processing unit (CPU), rather than replacing its general system-management role. (learn.microsoft.com)
Historical development
Graphics hardware developed from specialized rendering functions toward increasingly programmable processors. In 1999, NVIDIA marketed the GeForce 256 as the first GPU under its definition of a single-chip processor integrating transformation, lighting, triangle setup or clipping, and rendering. This was a product-specific definition, not evidence that graphics acceleration began in 1999. The integration of these functions moved more graphics work into dedicated hardware. (nvidia.com)
Programmable shading and general-purpose programming expanded the range of GPU applications. NVIDIA’s CUDA computing architecture emerged in 2006, allowing developers to use GPU execution resources for workloads beyond graphics. An influential example was the 2012 AlexNet study, whose convolutional neural network was trained on two NVIDIA GTX 580 GPUs. The paper demonstrated the practical value of GPU acceleration for large-scale visual recognition. (nvidianews.nvidia.com)
Architecture and execution
GPUs exploit parallel computing by maintaining many lightweight execution threads. Their organization favors data parallelism, in which similar operations are applied to numerous data elements. Large arrays and matrices often provide this structure, whereas a workload dominated by sequential dependencies offers fewer opportunities to use the hardware effectively. GPU acceleration therefore depends on the workload and its implementation, not simply on the presence of a GPU. (docs.nvidia.com)
NVIDIA describes its execution model as single instruction, multiple threads (SIMT). Threads retain individual state and control flow, but hardware schedules them in groups called warps. Other GPU architectures use related group-based execution arrangements. When threads within a group follow different conditional paths, branch divergence can reduce execution efficiency because the paths cannot always proceed simultaneously. (docs.nvidia.com)
Programmable units are supplemented by specialized engines. For example, some GPUs include hardware that accelerates ray tracing, particularly the traversal of spatial structures and ray–triangle intersection tests. Some also contain matrix-processing units, such as NVIDIA’s Tensor Cores, designed to accelerate operations important to deep learning. These units are distinct from both general shader execution resources and a standalone tensor processing unit. (nvidia.com)
Graphics rendering
In a conventional graphics pipeline, an application supplies geometric primitives, associated data, and rendering state. Programmable shaders process vertices and determine surface appearance. Rasterization identifies the image samples covered by geometric primitives; subsequent processing calculates their colors and combines results through depth, stencil, and blending operations. Not every application uses every available pipeline stage. (learn.microsoft.com)
Ray tracing follows a different approach, calculating intersections between rays and scene geometry to determine visibility and lighting effects. Rendering systems can combine rasterization with ray-traced shadows, reflections, or illumination. Dedicated ray-tracing hardware accelerates selected parts of this process; it does not eliminate the need for programmable shading, scene preparation, or image reconstruction. (nvidia.com)
Integration and memory
An integrated GPU is incorporated into the processor or its package, while a discrete GPU is a separate processor. Integrated designs commonly access system memory; discrete designs commonly have dedicated device memory. A graphics card is the board-level assembly containing a GPU and supporting components, not the processor itself. Systems can use integrated and discrete graphics together, assigning work according to their capabilities. (cdrdv2-public.intel.com)
Within a GPU, memory is organized into spaces with different access characteristics and scopes. These include registers, caches, device-wide memory, and fast local storage available to cooperating thread groups. CUDA calls the latter shared memory. Efficient programs reuse data in fast storage and arrange accesses so that requests from neighboring threads can be combined into fewer memory transactions. (docs.nvidia.com)
Memory capacity determines how much data can reside on the device, while bandwidth determines how quickly data can be supplied to its execution units. Moving data between CPU and GPU memory can impose substantial overhead. Some integrated systems share physical memory between host and device, although programming interfaces may still distinguish their logical memory spaces. (docs.nvidia.com)
Programming and computational applications
General-purpose computing on GPUs uses programmable execution resources for nongraphics tasks. Applications commonly divide work between host code running on a CPU and parallel kernels running on an accelerator. CUDA provides NVIDIA’s programming environment, while OpenCL is an open standard supporting GPUs and other heterogeneous processors. Libraries and higher-level frameworks can invoke kernels without requiring application authors to implement every low-level operation. (khronos.org)
GPU applications include scientific software, image processing, and neural-network training and inference. Their suitability often derives from repeated arithmetic over large datasets. AlexNet, for example, used optimized GPU implementations of convolution and other training operations. Its two-GPU arrangement also addressed the memory limits of an individual device. (khronos.org)
Performance constraints
GPU performance depends on available parallelism, memory access patterns, instruction throughput, and communication overhead. Small tasks may not provide enough work to offset kernel-launch and transfer costs. Irregular accesses, divergent branches, or excessive storage requirements can leave execution resources underused. Consequently, peak arithmetic throughput alone does not establish application performance: measurements must account for data movement and the portions of the program that remain outside the accelerated kernels. (docs.nvidia.com)