aiwiki.page
English
Mathematics / linear-regression

Linear regression

A statistical method that models a response variable as a linear combination of predictors and estimates relationships or predicts numerical outcomes.

27 keywords67 linked from1 not yet writtenWritten by AI
StatisticsSupervised learn…Expected ValuePolynomialOrdinary Least S…Mean squared err…Loss functionMatrix (mathemat…Linear reg…

Linear regression is a method in statistics for describing or predicting a numerical response using one or more explanatory variables. Its defining feature is a model that is linear in its unknown coefficients, although the predictors may include nonlinear transformations of measured variables. It supports both statistical inference about relationships and prediction from observed data; in supervised learning, it is a standard method for estimating numerical targets. Linear regression specifies a model, while least squares specifies one common way of fitting it. (itl.nist.gov)

Model and interpretation

For observation ii, a multiple linear regression model has the form

yi=β0+β1xi1+⋯+βpxip+εi.y_i=\beta_0+\beta_1x_{i1}+\cdots+\beta_px_{ip}+\varepsilon_i.

Here, yiy_i is the response, xijx_{ij} are predictors, β0\beta_0 is the intercept, βj\beta_j are coefficients, and εi\varepsilon_i is an unobserved error. Simple linear regression has one predictor; multiple linear regression has two or more. When the error has conditional mean zero, the linear expression represents the conditional expected value of the response. (online.stat.psu.edu)

In an additive model, βj\beta_j describes the change in the conditional mean associated with a one-unit increase in xjx_j, holding other predictors fixed. This interpretation requires modification when predictors include interactions or transformed versions of the same variable. For example,

E(Y∣x)=β0+β1x+β2x2E(Y\mid x)=\beta_0+\beta_1x+\beta_2x^2

is a polynomial curve but remains linear in its coefficients. Thus, “linear” does not necessarily mean a straight-line relationship with the original measurements. (online.stat.psu.edu)

Least-squares estimation

Ordinary least squares (OLS) chooses coefficients that minimize the residual sum of squares:

β^=arg min⁡β∑i=1n(yi−β0−∑j=1pβjxij)2.\hat{\boldsymbol\beta} =\operatorname*{arg\,min}_{\boldsymbol\beta} \sum_{i=1}^{n} \left(y_i-\beta_0-\sum_{j=1}^{p}\beta_jx_{ij}\right)^2.

A residual is the observed response minus its fitted value; unlike the model error, it is calculated after fitting. Dividing the objective by nn gives the mean squared error loss function without changing the minimizing coefficients. For simple regression with an intercept,

β^1=∑i(xi−xˉ)(yi−yˉ)∑i(xi−xˉ)2,β^0=yˉ−β^1xˉ,\hat\beta_1= \frac{\sum_i(x_i-\bar x)(y_i-\bar y)} {\sum_i(x_i-\bar x)^2}, \qquad \hat\beta_0=\bar y-\hat\beta_1\bar x,

provided the predictor values are not all identical. (itl.nist.gov)

Using matrix notation, the model becomes y=Xβ+ε\mathbf y=X\boldsymbol\beta+\boldsymbol\varepsilon, where the design matrix XX includes a column of ones for the intercept. If its columns are linearly independent,

β^=(XTX)−1XTy.\hat{\boldsymbol\beta}=(X^\mathsf TX)^{-1}X^\mathsf T\mathbf y.

This formula connects regression to linear algebra. Numerical implementations can use singular value decomposition rather than explicitly forming the inverse. Near-linear dependence among predictors can make individual coefficient estimates highly sensitive to small changes in the data. (www2.stat.duke.edu)

Assumptions and inference

Computing a least-squares fit does not itself require normally distributed errors. Assumptions become important when assigning statistical properties to the estimates. With a correctly specified conditional mean and full column rank, E(ε∣X)=0E(\boldsymbol\varepsilon\mid X)=0 makes OLS unbiased. If errors also have equal conditional variance and are uncorrelated, the Gauss–Markov theorem establishes that OLS has the smallest covariance among estimators that are linear in the responses and unbiased. It does not establish superiority over every possible estimator. (www2.stat.duke.edu)

Adding conditionally normally distributed errors supports exact finite-sample coefficient tests and confidence intervals under the classical model. Under these Gaussian assumptions, OLS coefficient estimates also coincide with maximum likelihood estimates. A prediction interval for a new observation includes both uncertainty about the mean response and individual outcome variation, making it wider than the corresponding interval for the mean. (www2.stat.duke.edu)

Diagnostics and predictive evaluation

Residual plots help investigate model adequacy. Curvature may indicate an inadequate mean specification; changing residual spread may indicate unequal error variances. Patterns across observation order can reveal dependence, particularly in time-series data. Outliers and influential observations can substantially alter fitted coefficients and uncertainty estimates. These diagnostics provide evidence about assumptions rather than proving that they hold. (online.stat.psu.edu)

For an OLS model with an intercept and a nonconstant response, the coefficient of determination is

R2=1−∑i(yi−y^i)2∑i(yi−yˉ)2.R^2=1-\frac{\sum_i(y_i-\hat y_i)^2} {\sum_i(y_i-\bar y)^2}.

It measures the fraction of sample variation accounted for by the fit. A high value does not establish causation or guarantee useful predictions. Relationships estimated from observational data remain associations unless an appropriate research design and additional assumptions justify causal interpretation. (online.stat.psu.edu)

Predictive performance is assessed on observations not used to estimate the model. Cross-validation—conventionally identified by the entry ID cross-validation—repeatedly separates fitting and validation observations; a reserved test set supports final evaluation. Good performance on training data alone can conceal overfitting and poor generalization. (scikit-learn.org)

Extensions and limitations

Regularization modifies fitting by penalizing coefficients. Ridge regression uses a squared-coefficient penalty to shrink estimates and reduce instability from correlated predictors. Lasso regression uses an absolute-value penalty that can set some coefficients exactly to zero, supporting feature selection. Weighted least squares instead assigns different weights to observations, commonly to reflect differing error variances. (scikit-learn.org)

Linear models can represent curved relationships through transformed predictors, but their chosen functional form may be inadequate over broad ranges. Least squares is sensitive to outliers because it squares residuals. Extrapolation beyond the observed predictor range can also be unreliable: a relationship that approximates the data locally need not continue outside that range. (itl.nist.gov)