Mean squared error (MSE) is a measure of error used in statistics and machine learning. It quantifies the average squared difference between an estimate or prediction and the value being estimated or predicted. In statistical estimation, the average is an expectation over repeated samples; in prediction assessment, it is commonly an arithmetic average over observations. MSE serves both as an evaluation measure and as a loss function for fitting models. Squaring prevents positive and negative errors from cancelling and gives larger errors disproportionately greater weight. (stat.berkeley.edu)
Definitions and basic properties
For (n) observed values (y_i) and corresponding predictions (\hat y_i), empirical MSE is
[ \operatorname{MSE} =\frac{1}{n}\sum_{i=1}^{n}(y_i-\hat y_i)^2. ]
It is nonnegative and equals zero precisely when every prediction matches its observed value. For example, errors of (1,-2,3) yield an MSE of ((1+4+9)/3=14/3). Nonnegative observation weights (w_i), with a positive total, give the weighted version
[ \operatorname{MSE}_w =\frac{\sum_i w_i(y_i-\hat y_i)^2}{\sum_i w_i}. ]
For multiple response variables, errors may be reported separately or aggregated using specified output weights. (scikit-learn.org)
In estimation theory, an estimator (\hat\theta) of a fixed parameter (\theta) has MSE
[ \operatorname{MSE}{\theta}(\hat\theta) =\mathbb E{\theta}[(\hat\theta-\theta)^2]. ]
Here, the expectation concerns the estimator’s sampling distribution, not an average over coordinates or observations in one realized dataset. These definitions express the same squared-error principle but refer to different averaging operations. (stat.berkeley.edu)
MSE has squared units: errors measured in metres produce MSE in square metres. Multiplying all errors by (c) multiplies MSE by (c^2). Its square root, root mean squared error (RMSE), restores the original units. For a single fixed evaluation dataset and aggregation rule, MSE and RMSE rank models identically because the square root is increasing. (scikit-learn.org)
Bias, variance, and prediction risk
Provided the relevant second moment exists, estimator MSE decomposes as
[ \operatorname{MSE}{\theta}(\hat\theta) =\operatorname{Var}{\theta}(\hat\theta) +\bigl(\mathbb E_{\theta}[\hat\theta]-\theta\bigr)^2. ]
The first component is variance, describing sampling variability; the second is squared estimator bias, describing systematic displacement from the parameter. An unbiased estimator therefore has MSE equal to its variance. Unbiasedness alone does not guarantee the smallest MSE: a biased estimator can perform better if its variance reduction outweighs its squared bias. (stat.berkeley.edu)
Prediction error has a related decomposition. Suppose (Y=f(X)+\varepsilon), with (\mathbb E[\varepsilon\mid X]=0). At a fixed input (x), averaging over training samples and an independent new response gives
[ \mathbb E[(Y-\hat f(x))^2\mid X=x] =\operatorname{Bias}(\hat f(x))^2 +\operatorname{Var}(\hat f(x)) +\operatorname{Var}(\varepsilon\mid X=x). ]
The final term is irreducible response variability under the specified information set. This distinction underlies the bias–variance tradeoff: reducing training error need not reduce expected prediction error. (nlp.stanford.edu)
Optimal prediction under squared loss
For a random variable (Y) with finite second moment, the constant (a) minimizing (\mathbb E[(Y-a)^2]) is (\mathbb E[Y]). With predictors (X), the corresponding optimal predictor is the conditional expectation
[ f^*(x)=\mathbb E[Y\mid X=x]. ]
This follows from the identity
[ \mathbb E[(Y-a)^2\mid X=x] =\operatorname{Var}(Y\mid X=x) +\bigl(\mathbb E[Y\mid X=x]-a\bigr)^2. ]
Thus squared loss targets the conditional mean rather than the median or most probable response. Mean absolute error, by contrast, targets a conditional median. The choice of loss consequently determines which aspect of the response distribution a prediction represents. (scikit-learn.org)
Model fitting and optimization
In supervised learning, minimizing empirical MSE on training data is an instance of empirical risk minimization. For linear regression, predictions are linear in the fitted coefficients, and minimizing MSE yields the same coefficients as ordinary least squares. Dividing the sum of squared errors by (n), or additionally by two, does not change its minimizers. (cs229.stanford.edu)
For a scalar prediction, the squared-loss derivative is (2(\hat y-y)). This smooth expression supports gradient descent and stochastic gradient descent. Although squared loss is a convex function of the prediction, composing it with a nonlinear parameterized model does not generally produce a convex function of its parameters. (cs229.stanford.edu)
Under independent, zero-mean errors following a normal distribution with common variance, minimizing squared error is equivalent to maximum likelihood estimation of the regression mean parameters. Gaussian assumptions are unnecessary for calculating MSE; they provide this particular probabilistic interpretation. (cs229.stanford.edu)
Interpretation and limitations
MSE is sensitive to unusually large errors. Doubling an error quadruples its contribution, so a few extreme observations can dominate the result. Absolute-error loss grows linearly, while Huber loss combines quadratic behaviour near zero with linear behaviour for sufficiently large errors. These losses encode different penalties rather than interchangeable versions of one criterion. (scikit-learn.org)
Low training MSE does not establish generalization. Overfitting can produce small training errors alongside substantially larger errors on unseen observations. Evaluation on a reserved test set or through cross-validation separates fitting performance from predictive assessment; repeatedly selecting models using test results compromises that separation. (scikit-learn.org)
Finally, regression terminology requires care. In classical regression and analysis of variance, “mean square error” often denotes the residual sum of squares divided by residual degrees of freedom. For simple linear regression with an intercept, this divisor is (n-2), not (n). That quantity estimates the model’s error variance and is distinct from the empirical prediction MSE defined above. (online.stat.psu.edu)