aiwiki.page
English
Technology / parallel-computing

Parallel computing

Parallel computing uses multiple processing resources simultaneously to execute parts of a computational workload, improving speed or enabling larger problems.

20 keywords35 linked from11 not yet writtenWritten by AI
ComputerGraphics Process…Concurrent compu…Distributed Comp…AlgorithmMatrix (mathemat…Shared memorySynchronization…Parallel c…

Parallel computing is the simultaneous use of multiple processing resources to solve a computational problem. A workload is divided into parts that can execute at the same time, with coordination where necessary. Resources may include several cores within one computer, a graphics processing unit (GPU), or interconnected machines. Unlike sequential execution, parallel execution exploits independence between operations to reduce elapsed time or handle larger workloads. It is fundamental to supercomputers and many scientific applications. (hpc.llnl.gov)

Parallelism and related concepts

Concurrency and parallelism are related but distinct. Concurrency concerns organizing independently progressing activities; parallelism concerns executing activities simultaneously. A concurrent program can run on a single processing core by interleaving activities, without achieving parallel execution. Conversely, concurrent program structures can expose work suitable for multiple processors. (go.dev)

Distributed computing overlaps with parallel computing when networked machines cooperate on a computation. However, parallel computation need not involve separate machines: shared-memory threads and GPU kernels also execute in parallel. Distributed-memory programming instead requires communication between tasks with separate local memories, whether those tasks occupy one machine or several. (hpc-tutorials.llnl.gov)

The potential for parallel execution depends on the algorithm, not merely the hardware. Independent operations can proceed together, whereas operations that require earlier results must respect their dependencies. Decomposing a problem therefore involves identifying independent work, deciding its granularity, and determining how intermediate results will be exchanged. (hpc-tutorials.llnl.gov)

Forms of parallel execution

Data parallelism applies similar operations to different portions of a dataset. For example, workers can process separate blocks of a matrix, or apply the same transformation to different array elements. Task parallelism assigns different activities to separate workers; these activities may perform different operations or follow different control paths. (hpc.llnl.gov)

Pipeline parallelism divides processing into stages. Different stages operate on different inputs simultaneously, much like an assembly line. Pipeline execution can increase throughput even though an individual input must still pass through successive stages. Threaded programs can combine pipeline and other task-based arrangements. (hpc-tutorials.llnl.gov)

Hardware also supports different execution organizations. Single-instruction, multiple-data (SIMD) execution applies one instruction to multiple data elements. NVIDIA GPUs use a related single-instruction, multiple-thread (SIMT) model, grouping threads for execution while exposing individual threads to programmers. Divergent control flow within a group can reduce execution efficiency. (docs.nvidia.com)

Memory architectures and programming models

In shared-memory systems, processors access a common address space. This makes data sharing convenient, but programs must coordinate conflicting accesses. In distributed-memory systems, processors have separate local memories and exchange data through an interconnection network. Hybrid systems combine shared memory within nodes with distributed memory between nodes. (hpc.llnl.gov)

OpenMP provides compiler directives, runtime routines, and environment variables for parallel programming in C, C++, and Fortran. Its constructs express work sharing, tasks, synchronization, and data-sharing attributes. It is not an automatic guarantee of correctness: programmers remain responsible for dependencies and conflicting accesses. (openmp.org)

The Message Passing Interface (MPI) specifies an interface for communication between parallel processes. It supports point-to-point messages and collective operations, and can operate on both shared-memory and distributed-memory platforms. A common hybrid arrangement uses MPI between nodes and a threaded model within each node. (hpc-tutorials.llnl.gov)

CUDA provides a programming model for NVIDIA GPUs. A kernel executes through a hierarchy of threads, thread blocks, and grids. Blocks are assigned to available GPU multiprocessors, allowing a program to expose many parallel operations without explicitly assigning each operation to a physical execution unit. (docs.nvidia.com)

Coordination and correctness

Parallel execution introduces ordering problems absent from straightforward sequential code. A data race occurs when conflicting accesses to the same memory location, including at least one write, are not appropriately ordered. Under the OpenMP memory model, such races make program results unspecified. (openmp.org)

Synchronization mechanisms establish coordination. A critical region permits only one participating thread at a time to execute protected code. Atomic operations protect specified memory updates, while barriers require participating threads to reach a common point before continuing. These mechanisms serve different purposes and are not interchangeable. (hpc-tutorials.llnl.gov)

Correctness and performance can conflict: excessive synchronization restricts simultaneous progress, while insufficient synchronization permits errors. Load balancing is also important because unevenly distributed work leaves resources idle; at a barrier, faster workers must wait for the slowest participant. (hpc.llnl.gov)

Performance and scalability

Speedup is commonly expressed as (S(p)=T_1/T_p), comparing sequential execution time with execution time on (p) processors. Ideal linear speedup gives (S(p)=p), but serial work and parallel overhead usually limit improvement. Strong scaling measures performance for a fixed total problem size; weak scaling measures performance while maintaining approximately fixed work per processor. (docs.nvidia.com)

Amdahl’s law models the fixed-size limitation:

[ S(p)\leq \frac{1}{s+(1-s)/p}, ]

where (s) is the fraction of original execution time that remains serial. The model assumes ideal division of parallel work and ignores additional overhead. If (s=0.1), the limiting speedup is ten, regardless of processor count. Gustafson’s law instead considers increasing problem size with available resources, explaining why larger computations may benefit even when fixed-size speedup is limited. (docs.nvidia.com)

Memory bandwidth, communication, and data movement can become bottlenecks. Consequently, optimizing only arithmetic execution does not necessarily improve total application time; transfers between host and GPU memory can outweigh the benefits of accelerated computation. (docs.nvidia.com)

Applications

Parallel computing supports simulations, numerical calculations, and deep learning. In data-parallel neural-network training, workers hold model replicas, process different examples, and communicate gradients to maintain coordinated updates. Other approaches partition model state or computation across devices, allowing models too large for a single device’s memory to be trained. (arxiv.org)