Ordinary least squares (OLS) is an estimation method in statistics that determines the coefficients of a linear regression model by minimizing the sum of squared residuals—the differences between observed responses and fitted values. “Ordinary” indicates that every observation receives the same weight in this criterion. OLS is a method for fitting a model, rather than a model itself; its statistical interpretation depends on assumptions about how the observations were generated. (itl.nist.gov)
Model and objective
For observations and coefficients, the regression model is written
where is the response vector, is an design matrix, is the unknown coefficient vector, and contains unobserved errors. If an intercept is included, one column of consists of ones, and the intercept counts among the coefficients. OLS estimates through the optimization problem
The squared-error loss function gives increasingly large penalties to large residuals. Dividing the objective by produces the mean squared error without changing its minimizers. (statlect.com)
“Linear” means linear in the unknown coefficients, not necessarily in the original predictor variables. A model such as is therefore an OLS-compatible polynomial regression. Fixed transformations and interactions can likewise become columns of the design matrix. (itl.nist.gov)
Algebraic solution and geometry
Setting the gradient of the objective to zero gives the normal equations,
When has full column rank, the solution is unique:
For a single predictor and an intercept, this reduces to
Thus the fitted line passes through the point of sample means. A unique slope requires that the predictor values are not all identical. (statlect.com)
In linear algebra, least squares has a geometric interpretation: the fitted vector is the orthogonal projection of onto the subspace spanned by the columns of . The residual vector satisfies . Consequently, residuals sum to zero when an intercept is present. If columns are linearly dependent, coefficient vectors need not be unique, although the fitted vector remains unique. A minimum-norm coefficient solution can be obtained through singular value decomposition. (netlib.org)
Statistical assumptions and properties
The least-squares calculation does not require a particular error distribution. Statistical guarantees require additional conditions. Under the correctly specified model, full column rank, and the zero conditional expectation condition
OLS is conditionally unbiased: . This condition is stronger than requiring the errors merely to have an unconditional mean of zero. (qed.econ.queensu.ca)
If, additionally,
errors have equal conditional variance and zero pairwise conditional covariance. The Gauss–Markov theorem then establishes that OLS is the best linear unbiased estimator: every linear combination of its coefficients has minimum variance among estimators linear in and unbiased under these assumptions. Its conditional covariance matrix is
Neither normality nor independence is required for this theorem; the stated covariance condition is sufficient. (qed.econ.queensu.ca)
With conditionally Gaussian errors, OLS also coincides with maximum likelihood estimation of the regression coefficients. Under suitable sampling, moment, and identification conditions, OLS is consistent and asymptotically normal even without Gaussian errors. (qed.econ.queensu.ca)
Inference and diagnostics
Under the classical assumptions, for , an unbiased estimate of error variance is
The denominator represents residual degrees of freedom. Estimated coefficient standard errors are the square roots of the diagonal entries of . Under Gaussian errors, these support exact finite-sample -based confidence intervals and tests of individual coefficients. (qed.econ.queensu.ca)
Unequal error variances or correlated errors can invalidate conventional standard errors without necessarily making the coefficient estimator biased. Appropriate robust covariance estimators address inference under specified departures from the classical covariance assumptions; they do not repair a misspecified conditional mean or endogeneity. Residual plots can reveal curvature, changing variability, and unusual observations, but cannot establish every statistical assumption. (statlect.com)
Computation and related methods
The inverse formula describes the estimator mathematically, but software commonly solves least-squares problems using QR factorization or singular value decomposition. These approaches avoid explicitly forming the inverse and are particularly important when predictors are nearly linearly dependent. LAPACK provides routines for full-rank, rank-deficient, and minimum-norm least-squares problems. (netlib.org)
OLS is sensitive to outliers and may extrapolate poorly beyond the observed predictor range. Weighted least squares modifies the objective by assigning observation-specific weights; inverse-variance weights are relevant when error variances differ and errors are uncorrelated. Regularization instead modifies estimation through penalties or constraints. For example, ridge regression penalizes squared coefficient magnitudes, potentially reducing estimation variance at the cost of introducing bias. (itl.nist.gov)