Floating-point arithmetic is a method of performing arithmetic on numbers represented by a sign, a finite-precision significand, and an exponent. It enables a computer to handle values spanning a wide range of magnitudes without storing every intervening digit. Floating-point values form a finite set of representable numbers rather than the continuous system of real numbers. Calculations therefore involve rounding whenever an exact result cannot be represented. The principal standard, IEEE 754, specifies binary and decimal formats, operations, rounding behavior, and exceptional conditions; implementations may use hardware, software, or both. (standards.ieee.org)
Representation and formats
Floating-point representation resembles scientific notation. A finite value can be expressed as
where determines the sign, is the significand, is the radix, and is the exponent. The significand contains the significant digits; the exponent determines their scale. Unlike fixed-point arithmetic, the position of the radix point is not fixed relative to the stored digits. Precision and range are separate properties: significand length limits the detail retained, while exponent range limits the magnitudes available. (docs.oracle.com)
Common IEEE binary formats use base two and encode fields in bits. For normal numbers, the leading binary significand digit is implicitly 1, providing an additional precision bit without storing it explicitly. (docs.oracle.com)
| Format | Total bits | Exponent bits | Significand precision |
|---|---|---|---|
| Binary32, commonly single precision | 32 | 8 | 24 bits |
| Binary64, commonly double precision | 64 | 11 | 53 bits |
Binary32 provides roughly seven significant decimal digits, and binary64 roughly sixteen. Representable values are not uniformly spaced: within the normal range, absolute spacing generally increases with magnitude. Consequently, a small increment may disappear when added to a sufficiently large value. (docs.oracle.com)
Rounding and arithmetic operations
IEEE 754 defines the behavior of basic operations such as addition, subtraction, multiplication, division, and square root. A correctly rounded operation returns the value obtained by computing the exact mathematical result and then rounding it to the destination format. This describes the required result, not necessarily the internal procedure used by hardware. (docs.nvidia.com)
The default binary rounding direction is nearest, with exact halfway cases resolved toward the result whose least significant significand digit is even. Directed modes round toward positive infinity, negative infinity, or zero. A fused multiply-add computes with one final rounding, rather than rounding the product before adding . It can therefore produce a different, more accurate result than separate multiplication and addition. (docs.nvidia.com)
Representation error arises before arithmetic begins. The decimal fraction , for example, has an infinite repeating binary expansion. Under usual binary64 evaluation, 0.1 + 0.2 produces a value commonly displayed as 0.30000000000000004. This reflects approximated operands and a rounded sum, not an arbitrary malfunction. Decimal arithmetic can represent these particular inputs exactly, but finite decimal precision still requires rounding for results such as . (docs.python.org)
Special values and exceptions
IEEE formats include positive and negative zero, positive and negative infinity, and NaN values. Signed zeros compare equal in ordinary numerical comparisons but retain distinct signs that can affect subsequent operations. NaNs represent results of invalid operations, such as zero divided by zero; ordinary equality comparisons involving a NaN are false, including comparison with itself. (docs.oracle.com)
Subnormal numbers occupy the interval between zero and the smallest positive normal number. They retain progressively fewer significant digits as their magnitude decreases, enabling gradual underflow rather than an abrupt jump to zero. IEEE arithmetic distinguishes five exception conditions: invalid operation, division by zero, overflow, underflow, and inexact result. Default handling generally supplies a result and records a status flag, rather than necessarily terminating execution. Programming-language environments may impose additional behavior. (docs.oracle.com)
Numerical error and stability
Round-off error is often measured through relative error or a unit in the last place, abbreviated ulp. Machine epsilon commonly denotes the gap between 1 and the next larger representable value, although conventions differ. These measures describe arithmetic precision, not a universal bound on the accuracy of an entire calculation. (docs.oracle.com)
Floating-point addition is not generally associative: can differ from . Subtracting nearly equal approximations can cause catastrophic cancellation, exposing earlier errors in the few digits that remain. The subtraction itself may nevertheless be exact for the stored operands. (docs.nvidia.com)
Numerical stability concerns how an algorithm propagates errors. It differs from the conditioning of a problem, which measures sensitivity to changes in input. Pairwise summation and compensated summation can reduce error compared with straightforward accumulation. These distinctions are central to numerical linear algebra and other numerical computations. (docs.oracle.com)
Parallel computation and mixed precision
In parallel computing, changing the reduction tree changes the order of additions and may alter the final bits. Compiler transformations, intermediate precision, and the use of fused operations also affect numerical reproducibility. Identical mathematical formulas do not alone guarantee identical floating-point results across execution environments. (docs.nvidia.com)
Mixed-precision arithmetic combines formats within one computation. In machine learning, low-precision multiplication may be paired with higher-precision accumulation to reduce storage and computational costs. Bfloat16 uses 16 bits with an eight-bit exponent and eight bits of significand precision, including the implicit leading bit. It retains approximately binary32’s normal exponent range while representing values much more coarsely. (cloud.google.com)