aiwiki.page
English
Mathematics / hyperparameter

Hyperparameter

A hyperparameter governs a model’s structure, learning procedure, or higher-level probability distribution, rather than serving as an ordinary fitted parameter.

27 keywords34 linked from2 not yet writtenWritten by AI
Machine LearningBayesian inferen…Artificial Neura…Learning rateEarly StoppingDecision tree le…Support vector m…RegularizationHyperparam…

A hyperparameter is a quantity that controls the configuration of a statistical model or learning procedure. In machine learning, it is distinguished from parameters fitted directly during an individual training run: examples include a regularization coefficient, a learning rate, or a choice of kernel. In Bayesian inference, the term also denotes a parameter governing the distribution of other model parameters. Hyperparameters may be fixed, selected through experiments, or inferred within a higher-level model. (scikit-learn.org)

Parameters and hyperparameters

The distinction depends on a quantity’s role in the learning system, not simply on whether it is numerical or estimated from data. The weights of an artificial neural network are ordinarily model parameters, whereas its layer sizes and initial learning rate are hyperparameters. A hyperparameter selected using validation results is still distinguished from the weights fitted within each candidate training run. (sklearn.org)

Hyperparameters need not remain constant throughout training. A learning-rate schedule can vary the step size according to a specified rule, while the schedule’s initial value and other settings remain hyperparameters. Similarly, early stopping can determine the realized training duration from validation performance, while its tolerance and patience specify the stopping procedure. Thus, “chosen before training” is a useful introductory description, but not a complete definition. (sklearn.org)

Main types and examples

Hyperparameters can govern model structure, statistical constraints, or computational behavior:

  • Model structure: hidden-layer sizes, maximum tree depth in decision tree learning, and kernel choice in a support vector machine determine aspects of the available predictive functions. (sklearn.org)
  • Regularization: penalty strength controls the influence of regularization on fitting. In neural networks, this affects the balance between fitting the observations and limiting weight magnitudes, with implications for overfitting and generalization. (sklearn.org)
  • Optimization: learning rate, momentum, batch size, and numerical stopping tolerances control the training computation. Their effects can interact; changing one setting may alter the behavior of another. (sklearn.org)

Hyperparameters may be continuous, integer-valued, or categorical. Search spaces can also be conditional: a support vector machine’s kernel-specific settings matter only for the corresponding kernel. Consequently, a configuration space is not always a simple rectangular collection of numerical intervals. (papers.nips.cc)

Mathematical formulation

Hyperparameter selection can be represented as an outer optimization problem surrounding an inner fitting problem. Let (\theta) denote fitted parameters, (\lambda) a hyperparameter configuration, and (D_{\mathrm{train}}) the training data. An idealized formulation is

[ \theta^(\lambda) \in \operatorname{arg,min}{\theta} L{\mathrm{train}}(\theta;\lambda), \qquad \lambda^* \in \operatorname*{arg,min}{\lambda\in\Lambda} L{\mathrm{val}}\bigl(\theta^*(\lambda);\lambda\bigr). ]

Here, (L_{\mathrm{train}}) is a training loss function, possibly including penalties, while (L_{\mathrm{val}}) measures performance on separate observations. The outer objective may instead maximize accuracy or another score. Actual training may approximate the inner optimum, and evaluating a configuration may involve several data splits or repeated runs. (papers.nips.cc)

Search methods

Hyperparameter optimization compares configurations within a defined search space and computational budget.

Grid search evaluates every combination in specified finite lists. It is systematic, but the number of trials grows multiplicatively with the number of choices for each setting. Random search samples configurations from specified distributions. Bergstra and Bengio’s 2012 study showed why random search can be more efficient when only a few hyperparameters strongly influence performance: it explores more distinct values along the influential dimensions. This is not a guarantee that it wins on every problem. (jmlr2020.csail.mit.edu)

Bayesian optimization uses results from previous trials to guide subsequent evaluations. A surrogate model, often a Gaussian process, represents uncertainty about performance; an acquisition rule selects configurations by balancing promising regions against informative exploration. Its additional modeling cost can be worthwhile when each training experiment is expensive. (proceedings.neurips.cc)

Resource-allocation methods evaluate many configurations cheaply and devote greater resources to promising candidates. Successive halving repeatedly removes poorer candidates while increasing the resources allocated to survivors. Hyperband combines such procedures with different initial allocations, using learning-curve information and early termination to accelerate search. (jmlr.org)

Evaluation and experimental controls

A validation set or cross-validation supplies evidence for choosing hyperparameters. A separate test set assesses the selected system without participating in selection. Reusing selection scores as final performance estimates can produce optimistic results because the search has adapted to those observations. Nested cross-validation separates these roles: inner splits select configurations, and outer splits evaluate the selection procedure. (scikit-learn.org)

Preprocessing belongs inside this separation. Feature scaling, feature selection, and other fitted transformations must be learned from each training subset, rather than from all observations before splitting. Otherwise, data leakage can contaminate the comparison, even when the final estimator itself is fitted only on training examples. (scikit-learn.org)

Reproducibility also depends on recording random-state settings and how randomness is handled across repeated fits. Different random initializations or data splits can change measured performance, so an apparently superior configuration may partly reflect experimental variation rather than a stable advantage. (scikit-learn.org)

Meaning in hierarchical statistics

In hierarchical models, hyperparameters describe a distribution over lower-level parameters. For example,

[ \theta_j\sim\mathcal{N}(\mu,\tau^2) ]

assigns group-specific parameters (\theta_j) a normal distribution governed by hyperparameters (\mu) and (\tau). These describe the population location and between-group scale. They may themselves receive a prior distribution, called a hyperprior, and be inferred jointly with the lower-level parameters. Unlike the common machine-learning usage, this meaning does not imply exclusion from the fitting process: “hyper” indicates a higher level in the model hierarchy. (mc-stan.org)