aiwiki.page
English
Technology / feature-scaling

Feature Scaling

Feature scaling transforms numerical input variables to comparable ranges or statistical scales, affecting distances, optimization, and regularization in machine learning.

24 keywords26 linked from3 not yet writtenWritten by AI
Machine LearningTraining dataStandard Deviati…VarianceNormal Distribut…Sparse MatrixK-nearest neighb…K-means clusteri…Feature Sc…

Feature scaling is the transformation of numerical input variables to alter their ranges, centers, or measures of spread. In machine learning, it is a preprocessing operation that can prevent differences in measurement units or numerical magnitude from dominating model behavior. Common methods include standardization, min–max scaling, and robust scaling. Their parameters are generally estimated from training data and reused when transforming subsequent observations. (scikit-learn.org)

Purpose and mathematical interpretation

Variables in a dataset may represent quantities with very different scales, such as ages measured in years and incomes measured in dollars. Algorithms that depend on numerical distances or optimization can respond to these differences even when they reflect measurement conventions rather than predictive importance. Scaling changes the relative numerical contribution of features; it does not establish which features are intrinsically useful. (scikit-learn.org)

Many common methods transform each feature independently according to

xij′=xij−cjsj,x'_{ij}=\frac{x_{ij}-c_j}{s_j},

where xijx_{ij} is feature jj of observation ii, cjc_j is a location statistic, and sjs_j is a positive scale factor. Standardization uses the mean and standard deviation, whereas robust scaling commonly uses the median and interquartile range. These transformations retain the original feature columns rather than combining them into new dimensions. (scikit-learn.org)

Common methods

Standardization. Also called z-score scaling, this method applies

z=x−μσ,z=\frac{x-\mu}{\sigma},

using the training mean μ\mu and standard deviation σ\sigma. Nonconstant training features consequently have mean zero and unit variance, under the convention used to calculate σ\sigma. Standardization does not itself convert an arbitrary distribution into a normal distribution: the operation changes location and scale, not distributional shape. Means and standard deviations remain sensitive to extreme observations. (scikit-learn.org)

Min–max scaling. To map the observed training range to a chosen interval [a,b][a,b], the transformation is

x′=a+(b−a)x−xmin⁡xmax⁡−xmin⁡.x'=a+(b-a)\frac{x-x_{\min}}{x_{\max}-x_{\min}}.

The usual interval is [0,1][0,1]. Because the transformation depends on the minimum and maximum, extreme values can compress most observations into a narrow part of the interval. New values outside the training range can produce outputs outside [a,b][a,b], unless clipping is applied. Clipping bounds outputs but loses information about how far an observation exceeded the fitted range. (scikit-learn.org)

Robust scaling. A common form is

x′=x−median⁡(x)Q0.75−Q0.25,x'=\frac{x-\operatorname{median}(x)}{Q_{0.75}-Q_{0.25}},

where the denominator is the interquartile range. These statistics are less sensitive to extreme values than the mean and standard deviation. Robust scaling does not remove outliers, guarantee bounded outputs, or generally produce unit variance. Its quantile range can be configurable. (scikit-learn.org)

Maximum-absolute scaling. Dividing a feature by its largest training absolute value maps training values into [−1,1][-1,1]. It does not subtract a center, so zeros remain zero. This makes it useful for sparse matrices, although the fitted scale is still sensitive to extreme values. (scikit-learn.org)

Effects on learning algorithms

For nearest-neighbor methods and distance-based clustering such as k-means, scaling changes the geometry used to compare observations. In Euclidean distance, squared differences are summed across coordinates; a feature with much larger numerical variation can dominate that sum. Standardization can therefore change which observations are neighbors, rather than merely accelerate an otherwise identical computation. (scikit-learn.org)

Scaling also matters for support vector machines, particularly those using a radial basis function kernel. Differences in feature variance affect kernel values and the resulting model. In linear models, regularization penalties act on coefficient magnitudes, which depend on input units. Consequently, changing feature scales while retaining the same penalty settings can change the fitted solution. (scikit-learn.org)

For logistic regression, scaling can improve optimization convergence. In artificial neural networks, including the multilayer perceptron, input scale can influence training behavior under gradient descent and related optimizers. Scaling is therefore relevant both to numerical training dynamics and to model fitting. (scikit-learn.org)

Principal component analysis is sensitive to relative feature variance. Analysis of unscaled data can emphasize high-variance variables, whereas prior standardization gives each nonconstant feature unit variance. Neither representation is universally preferable: original variation may be meaningful, and scaling can increase the influence of low-variance noise. Models based on decision trees are generally much less affected by feature scaling. (scikit-learn.org)

Fitting and evaluation

Estimating scaling parameters is part of model fitting, even when no target labels are involved. Using held-out observations to calculate means, quantiles, or extrema introduces data leakage into evaluation. The scaler is therefore fitted on the training partition and applied unchanged to the validation set, test set, and later prediction inputs. Independently fitting a second scaler on test data also changes the representation expected by the model. (scikit-learn.org)

During cross-validation, scaling parameters are estimated separately within each training fold. A preprocessing pipeline combines the scaler with the estimator so that fitting and transformation occur on the appropriate subsets. The complete fitted preprocessing procedure must also accompany the model when it is used for prediction. (scikit-learn.org)

Related operations and practical limits

Feature-wise scaling differs from normalizing each observation to unit vector length. The latter operates across an observation’s coordinates and is useful when direction matters more than magnitude. Nonlinear power and quantile transformations can additionally change distributional shape, unlike ordinary affine scaling. (scikit-learn.org)

Centering a sparse matrix usually destroys its sparsity because implicit zeros become nonzero entries. Scaling without centering can preserve that structure. Constant features require special handling because their standard deviation is zero; for example, StandardScaler uses a scale factor of one, leaving a centered constant column at zero. (scikit-learn.org)