The coefficient of determination, usually denoted and pronounced “R-squared,” is a measure of model fit in statistics. In linear regression fitted by ordinary least squares with an intercept, it represents the proportion of the observed response’s variation about its mean accounted for by the fitted model. More generally, it compares squared prediction error with the error of a baseline that predicts the response mean. The distinction matters: ordinary in-sample lies between zero and one, whereas an calculated for arbitrary or out-of-sample predictions can be negative. (online.stat.psu.edu)
Definition and calculation
For observations , corresponding predictions , and sample mean
define the total sum of squares
and the residual sum of squares
The usual centered definition is
provided . The denominator measures the error of predicting for every observation; the numerator measures the model’s error on those same observations. (online.stat.psu.edu)
Because both quantities have the same squared units, their ratio is dimensionless. For a fixed evaluation dataset, maximizing is equivalent to minimizing mean squared error, since is fixed. However, equal mean squared errors on datasets with different response variability need not produce equal values. These properties follow directly from the definition. (scikit-learn.org)
Interpretation and range
For a nonconstant response, the definition gives:
- : every prediction equals its observed value.
- : the model has the same total squared error as predicting the evaluation dataset’s mean.
- : the model has greater squared error than that mean baseline.
There is no finite lower bound: sufficiently poor predictions can make arbitrarily large. Thus, despite its notation, the generalized prediction score need not be the square of a real-valued quantity. (scikit-learn.org)
For an ordinary least-squares fit with an intercept, the constant-mean model is among the available models. Least squares cannot fit the training observations worse than that baseline, so . For example, if and , then : the fitted model accounts for 75% of the observed variation about the mean. This does not mean that 75% of individual predictions are correct. (online.stat.psu.edu)
Least-squares decomposition and geometry
In ordinary least squares with an intercept, the total sum of squares decomposes as
Consequently,
This identity gives the familiar “explained variation divided by total variation” interpretation. It is not a universal identity for arbitrary predictions. (online.stat.psu.edu)
Geometrically, least squares is an orthogonal projection onto the linear span of the model’s predictor columns. The residual vector is orthogonal to the fitted vector. Including an intercept also makes the residuals sum to zero, so the centered response decomposes into orthogonal fitted and residual components. The sum-of-squares identity is therefore an instance of the Pythagorean theorem in a finite-dimensional inner-product space. This geometric interpretation follows from the least-squares decomposition. (online.stat.psu.edu)
Relationship to correlation
For simple ordinary least-squares regression with one predictor and an intercept,
where is the Pearson correlation coefficient between predictor and response. Squaring removes the sign: equally strong positive and negative linear relationships have the same . The regression slope or correlation coefficient is needed to determine direction. (online.stat.psu.edu)
For multiple ordinary least-squares regression with an intercept, also equals the squared correlation between observed and fitted responses, provided the fitted responses are nonconstant. This equivalence does not generally hold for arbitrary predictions. A prediction series can be perfectly correlated with the observations yet have substantial squared error because it has an incorrect offset or scale. (stats.oarc.ucla.edu)
Adjusted R-squared
Adding predictors to a nested ordinary least-squares model cannot increase its training residual sum of squares. As a result, training cannot decrease, even when the extra predictors contribute little meaningful information. This makes unadjusted insufficient by itself for selecting model complexity. (online.stat.psu.edu)
Adjusted R-squared incorporates a degrees-of-freedom correction. For a full-rank model with observations, an intercept, and predictors,
assuming . Unlike ordinary training , it can decrease when predictors are added. It is useful in model comparison, but it is not literally a proportion of explained variation and does not replace evaluation on unseen data. (online.stat.psu.edu)
Predictive evaluation
In machine learning, can be computed on a held-out test set or within cross-validation. Such scores assess predictive generalization, rather than merely fit to the observations used for estimation. A high training score alone does not establish good out-of-sample performance and may accompany overfitting. (scikit-learn.org)
The baseline must be stated carefully. In the conventional test-set calculation, is the test response mean. A score that instead compares the model against predictions made using the training response mean has a different denominator and can give a different result. Consequently, the phrase “better than predicting the mean” is incomplete unless the relevant mean is identified. (scikit-learn.org)
Alternative conventions and related measures
Models without an intercept. Some software reports an uncentered coefficient,
Its baseline is zero rather than the response mean. Centered and uncentered values are therefore not directly interchangeable. A no-intercept model can also be evaluated using the centered definition, in which case its training score may be negative. (statsmodels.org)
Explained-variance score. A related measure uses
where denotes variance. Unlike the usual , it does not penalize a constant prediction offset. The two scores coincide when the residuals have mean zero. (scikit-learn.org)
Pseudo-R-squared. Logistic regression and other generalized linear models often use pseudo- measures, including McFadden’s, Cox–Snell’s, and Nagelkerke’s measures. These use different constructions, frequently involving a likelihood function, and generally cannot be interpreted as the ordinary least-squares proportion of explained variation. Comparisons require the same definition, outcome, and dataset. (stats.oarc.ucla.edu)
Limitations and undefined cases
A high does not establish causation, demonstrate that the chosen functional relationship is correct, or guarantee useful predictions. A low value does not by itself rule out a meaningful association. Residual patterns, uncertainty, and the application’s required prediction precision contain information that a single fit statistic cannot express. There is no universal threshold separating a “good” model from a “bad” one. (online.stat.psu.edu)
If all observed responses are identical, , so the usual mathematical definition is undefined. Software may substitute conventional finite scores; for example, scikit-learn’s documented default returns 1 for perfect predictions and 0 for imperfect predictions in this case. Such substitutions are implementation conventions, not consequences of the defining formula. The score is also not well-defined for a single observation. (scikit-learn.org)
References
- r2_score — scikit-learn 1.6.1 documentationscikit-learn.org
- 4. Metrics and scoring: quantifying the quality of predictions — scikit-learn documentationscikit-learn.org
- statsmodels.regression.linear_model — statsmodels documentationstatsmodels.org
- FAQ: What are pseudo R-squareds?stats.oarc.ucla.edu