The bias–variance tradeoff is a concept in statistics and machine learning describing the relationship between systematic prediction error and variability across training samples. A learning method may produce stable predictions that consistently miss the underlying relationship, or more accurate average predictions that change substantially when the data change. The objective is good generalization: minimizing expected error on unseen observations rather than either component individually. Under squared-error loss, prediction error has an exact decomposition into squared bias, variance, and irreducible noise. The familiar claim that greater model complexity lowers bias but raises variance is a useful pattern, not a universal law. (cs.cmu.edu)
Bias and variance
In supervised learning, an algorithm constructs a predictor from training data. Imagine repeating this process with many independently sampled datasets of the same size from the same population. At any fixed input, the fitted prediction becomes a random variable because it depends on the sampled dataset. (cs.cmu.edu)
Bias is the difference between the predictor’s average output over these datasets and the target value. Variance measures how widely its predictions fluctuate around that average. High bias indicates systematic inaccuracy; high variance indicates sensitivity to the particular training sample. A predictor can have low bias without being reliable on any individual dataset. Conversely, stable predictions need not be accurate. These definitions concern repeated sampling, not merely the spread of observed outcomes within one dataset. (scikit-learn.org)
Squared-error decomposition
Let denote a random training dataset and its fitted predictor. For a fresh response at input , write
where the conditional noise has mean zero and variance . Assume the fresh observation is independent of , and all relevant second moments exist. Define the mean prediction using the expected value over training datasets:
For the squared-error loss function, expected prediction error is
Thus, expected mean squared error equals squared bias plus variance plus noise. Adding and subtracting , then expanding the square, yields the identity: the centered cross terms have zero expectation. Averaging over the input distribution gives the corresponding population prediction error. When the target is itself rather than a new noisy response, the noise term is absent. (cs.cmu.edu)
Complexity and fitting
A restrictive model may exhibit underfitting because it cannot represent important structure. For example, a straight-line model fitted by linear regression cannot reproduce a strongly curved relationship. Allowing higher-degree polynomial terms increases flexibility, but can also make the fitted curve more sensitive to sampling fluctuations. Excessive sensitivity can contribute to overfitting, where excellent training performance does not translate into comparable predictive performance. (scikit-learn.org)
In the classical picture, increasing flexibility initially reduces squared bias enough to improve test error. Beyond an intermediate optimum, increasing variance outweighs further bias reduction, producing a U-shaped error curve. The optimum depends on the data distribution, sample size, noise, and fitting procedure; it is not an intrinsic property of a model family. (doi.org)
Controlling the tradeoff
Regularization restricts fitting through penalties or constraints. In ridge regression, penalizing squared coefficient magnitudes shrinks estimates toward zero. This introduces bias but can reduce variance and improve prediction, particularly when the unpenalized estimates are unstable. Stronger regularization does not necessarily improve performance: excessive shrinkage can discard useful structure. (github.com)
Another approach is averaging predictors. Bagging fits models to bootstrap resamples and combines their predictions. For unstable methods such as decision trees, averaging can substantially reduce variance. In a documented regression example, bagging slightly increases squared bias but reduces variance sufficiently to lower total prediction error. Its benefit depends on how differently the component models respond to the training data. (scikit-learn.org)
Empirical assessment
Bias and variance are defined over hypothetical repeated datasets, so they are not usually directly observable from a single fitted model. Cross-validation instead estimates predictive performance by repeatedly fitting on subsets and evaluating on held-out observations. It can compare model flexibility and regularization strength without separately estimating the decomposition’s components. A validation set supports model selection, while a separate test set supports evaluation after selection. (scikit-learn.org)
Learning curves compare training and validation performance as sample size changes. Poor performance on both can indicate high bias, whereas a substantial training–validation gap can indicate high variance. Such patterns are diagnostic rather than definitive measurements. Evaluation also requires preventing data leakage, including leakage from preprocessing fitted before the evaluation split. (scikit-learn.org)
Scope and modern qualifications
The displayed decomposition applies specifically to squared error; it cannot simply be substituted unchanged for other losses. More importantly, the identity does not require variance to increase monotonically with parameter count. In double descent, test error may rise near the capacity needed to interpolate training data and then fall as capacity increases further. Research has demonstrated this behavior in several model families, including neural networks. These findings qualify the classical U-shaped interpretation without invalidating the squared-error decomposition itself. (doi.org)