Feature selection is the process of choosing a subset of available input variables, or features, for use in a statistical or machine-learning model. It retains selected variables rather than combining them into new representations. Its objectives include improving prediction, reducing computation and measurement costs, and making models easier to interpret. Selection can remove irrelevant or redundant information, but a smaller feature set does not necessarily produce better predictions. (jmlr.org)
Scope and objectives
Feature selection is a form of dimensionality reduction, distinct from feature extraction. For example, principal component analysis constructs components from combinations of original variables, whereas selection retains particular original variables. It can also follow feature engineering: generated interactions or other derived variables may themselves become candidates for selection. (jmlr.org)
In supervised learning, selection uses a target variable to assess predictive usefulness. In unsupervised learning, criteria may instead concern variation or structure in the inputs; removing constant variables is a simple example. The appropriate criterion therefore depends on the task, not merely on a variable’s apparent importance. (scikit-learn.org)
An irrelevant variable supplies no useful information for the specified task; a redundant variable supplies information already available through other variables. These distinctions are contextual. Two weak individual predictors can be useful jointly, while several strongly predictive variables may encode nearly the same information. Removing inputs can reduce overfitting, but it can also discard useful signal. (jmlr.org)
Filter methods
Filter methods evaluate features using criteria separate from fitting the eventual predictive model. They commonly rank variables and retain a fixed number or those exceeding a threshold. Examples include correlation with the target, F-statistics based on analysis of variance, chi-square scores, and mutual information. Low-variance filtering does not require target labels. (scikit-learn.org)
Univariate filters assess each variable individually. They are computationally economical, but can overlook interactions and retain multiple redundant predictors. Linear association scores capture different relationships from mutual-information estimates, which can detect nonlinear dependence but require sufficient observations for reliable estimation. A ranking is not itself a selected subset: a cutoff or subset-size rule is also needed. (scikit-learn.org)
Wrapper methods
Wrapper methods assess candidate subsets through the performance of a particular learning procedure. Forward selection starts with an empty subset and adds variables; backward selection starts with a larger set and removes them. These are typically greedy algorithms, choosing a locally favorable change rather than evaluating every possible subset. Candidate performance is often estimated using cross-validation. (scikit-learn.org)
Recursive feature elimination repeatedly fits an estimator, ranks the remaining variables using its coefficients or importance scores, removes the least important, and refits. A cross-validated variant compares subset sizes. Because elimination depends on the estimator and earlier removals, its ranking is model-dependent rather than a universal ordering of usefulness. (scikit-learn.org)
Wrappers can account for how features work together within a model, but repeated fitting makes them expensive. Searching all subsets of (p) variables involves (2^p) possibilities, including the empty subset; practical methods generally search only part of this space. Evaluating many candidates also creates opportunities to overfit the selection criterion. (jmlr.org)
Embedded and model-based methods
Embedded methods incorporate selection into model fitting. Sparse regularization is a prominent example: lasso regression applies an (L_1) penalty that can set coefficients exactly to zero. Elastic net combines (L_1) and (L_2) penalties and can retain groups of correlated predictors more effectively than lasso alone. Conversely, ridge regression generally shrinks coefficients without making them exactly zero, so shrinkage does not automatically constitute selection. (scikit-learn.org)
Model-based selection can also retain variables whose fitted importance exceeds a threshold. Decision trees and random forests provide impurity-based scores, while linear models provide coefficients. Such scores reflect the fitted model and its assumptions. Their magnitudes are not interchangeable across estimator families, and coefficient-based comparisons can depend on variable scaling. (scikit-learn.org)
Evaluation and data leakage
Feature selection belongs inside the training procedure. Learning a subset from the complete dataset before evaluating a model exposes the selection process to evaluation observations. This data leakage can yield overly optimistic results—even when features and targets are unrelated. Selection rules must be fitted using training data, then applied unchanged to the corresponding evaluation data. A pipeline helps enforce this separation within cross-validation. (scikit-learn.org)
Subset size, selection thresholds, and regularization strength are model-selection choices. When these are tuned, nested cross-validation separates an inner selection process from an outer performance assessment. An independently reserved test set offers another final evaluation arrangement, provided it does not influence selection decisions. The object being evaluated is the complete selection-and-fitting procedure, not just a predictor trained on a previously chosen subset. (scikit-learn.org)
Stability and interpretation
Selection stability describes how consistently a method chooses features when observations or random seeds change. Small data perturbations can produce different subsets, so predictive accuracy alone does not characterize the reliability of feature preferences. Repeated resampling permits assessment of selection frequencies or agreement between subsets. (jmlr.org)
Correlated variables complicate interpretation. A model may obtain similar information from alternative predictors, making individual permutation feature importance scores small even when the group is useful. Different selected subsets may therefore support similar predictions. A feature’s inclusion establishes its usefulness under a particular dataset, model, and selection criterion; it does not by itself establish causation or a uniquely correct explanatory set. (scikit-learn.org)