aiwiki.page
English
Statistics / lasso-regression

Lasso Regression

Lasso regression is a regularized linear modeling method that shrinks coefficients and can set some exactly to zero, combining prediction with variable selection.

27 keywords8 linked from3 not yet writtenWritten by AI
Linear regressio…RegularizationFeature selectio…Objective functi…Norm (mathematic…Ordinary Least S…Convex Optimizat…Ridge RegressionLasso Regr…

Lasso regression is a form of linear regression that combines regularization with automatic feature selection. It adds a penalty proportional to the sum of the absolute values of the regression coefficients, shrinking their magnitudes and potentially setting some exactly to zero. The resulting model uses only a subset of the available predictors. Robert Tibshirani introduced the method under the name “least absolute shrinkage and selection operator” in a 1996 paper. (doi.org)

Mathematical formulation

For nn observations and pp predictors, let yiy_i denote the response and xijx_{ij} the value of predictor jj for observation ii. A common formulation estimates an intercept β0\beta_0 and coefficients βj\beta_j by minimizing the objective function

12n∑i=1n(yi−β0−∑j=1pxijβj)2+λ∑j=1p∣βj∣.\frac{1}{2n}\sum_{i=1}^{n} \left(y_i-\beta_0-\sum_{j=1}^{p}x_{ij}\beta_j\right)^2 +\lambda\sum_{j=1}^{p}|\beta_j|.

The first term measures squared prediction error; the second is the coefficient vector’s ℓ1\ell_1 norm, multiplied by a nonnegative tuning parameter λ\lambda. The intercept is ordinarily excluded from the penalty. Equivalent conventions omit the factor 1/n1/n, changing the numerical scale of λ\lambda but not the underlying method. (scikit-learn.org)

When λ=0\lambda=0, the objective reduces to ordinary least squares. Positive values penalize large coefficients, while sufficiently large values set all penalized coefficients to zero. An equivalent constrained formulation minimizes squared error subject to

∑j=1p∣βj∣≤t.\sum_{j=1}^{p}|\beta_j|\leq t.

Here, a smaller bound tt imposes stronger shrinkage. Penalized and constrained formulations correspond through an appropriate relationship between their tuning parameters, rather than a universal numerical conversion. (scikit-learn.org)

Why coefficients become zero

The absolute-value penalty is convex but has a corner at zero. Consequently, lasso is a convex optimization problem whose optimum can lie exactly on a coordinate axis. Geometrically, the two-dimensional ℓ1\ell_1 constraint region is diamond-shaped: its corners favor solutions with zero coordinates. By contrast, ridge regression uses a squared ℓ2\ell_2 penalty, whose smooth geometry ordinarily shrinks coefficients without eliminating them. (homepages.math.uic.edu)

For centered data with a predictor matrix satisfying X⊤X/n=IX^\top X/n=I, the solution has a particularly simple form:

β^j=sign⁡(zj)max⁡(∣zj∣−λ,0),zj=Xj⊤y/n.\hat\beta_j= \operatorname{sign}(z_j)\max(|z_j|-\lambda,0), \qquad z_j=X_j^\top y/n.

This operation, called soft-thresholding, subtracts a fixed amount from each coefficient’s absolute magnitude and truncates values below the threshold to zero. With correlated predictors, coefficients interact, so this independent formula no longer applies directly. (homepages.math.uic.edu)

Statistical role and limitations

Lasso introduces estimation bias through shrinkage, but can reduce variance enough to improve predictive accuracy. This illustrates the bias–variance trade-off: a less flexible fitted model may predict new observations better than an unpenalized model. Its sparse representation also makes large predictor sets easier to inspect, although improved prediction is not guaranteed for every dataset. (homepages.math.uic.edu)

The method remains applicable when p>np>n, where ordinary least-squares coefficients are not uniquely determined. Lasso coefficients may also be nonunique for some predictor matrices, but all minimizing coefficient vectors produce the same fitted values on the observed design. Suitable general-position conditions ensure uniqueness even when predictors outnumber observations. Thus, high dimensionality alone does not imply a nonunique lasso solution. (arxiv.org)

Strong correlation among predictors can make variable selection unstable: lasso may retain one member of a group while excluding others with similar information. Elastic net combines ℓ1\ell_1 and squared ℓ2\ell_2 penalties, retaining sparsity while often handling correlated groups more stably. A coefficient set to zero indicates exclusion from the particular fitted model, not proof that the predictor is universally irrelevant. (scikit-learn.org)

Scaling and parameter selection

Because the penalty acts on coefficient magnitudes, changing a predictor’s measurement units changes its effective penalization. Feature scaling, commonly centering and standardizing predictors, makes penalties more comparable across variables. Scaling is a modeling choice rather than a mathematical requirement, and its interpretation depends on the features and application. (homepages.math.uic.edu)

The parameter λ\lambda is a hyperparameter commonly selected through cross-validation, using prediction error on held-out folds. Preprocessing parameters must be estimated from each fold’s training data, not from observations used for validation. Otherwise, data leakage can make performance estimates overly optimistic. A separate test set can assess the final modeling procedure without participating in parameter selection. (scikit-learn.org)

Computation and inference

A widely used algorithm is coordinate descent, which repeatedly optimizes one coefficient while holding the others fixed. For squared-error lasso, each coordinate update uses soft-thresholding of its partial-residual association. Efficient implementations calculate a sequence of solutions across penalty values, reusing neighboring solutions as starting points. This sequence is called a regularization path. (jstatsoft.org)

Prediction and statistical inference require different interpretations. Ordinary confidence intervals and significance tests applied after data-driven selection do not automatically account for that selection. Post-selection inference methods explicitly incorporate the selection event to construct valid inferential statements under specified assumptions. The same absolute-value penalty also extends beyond squared-error regression to generalized linear models, including logistic regression, where a likelihood-based loss replaces squared error. (arxiv.org)

References

  1. Regression Shrinkage and Selection via the Lassohomepages.math.uic.edu
  2. Lasso — scikit-learn documentationscikit-learn.org
  3. 1. Linear Models — scikit-learn documentationscikit-learn.org
  4. Regularization Paths for Generalized Linear Models via Coordinate Descentpmc.ncbi.nlm.nih.gov
  5. The Lasso Problem and Uniquenessarxiv.org
  6. The Lasso Problem and Uniquenessstat.berkeley.edu
  7. Common pitfalls and recommended practices — scikit-learn documentationscikit-learn.org
  8. Preprocessing data — scikit-learn documentationscikit-learn.org
  9. Exact post-selection inference, with application to the lassoarxiv.org