Bootstrap sampling is a computational method in statistics that approximates the sampling distribution of an estimator by repeatedly generating samples from observed data or a fitted model. In its ordinary nonparametric form, it draws observations with replacement, usually producing samples of the same size as the original dataset. Recalculating a statistic on these samples provides estimates of uncertainty, including standard errors, bias, and confidence intervals. Bradley Efron introduced the bootstrap as a general statistical method in 1979. (blogs.helsinki.fi)
Principle and mathematical formulation
Suppose are observations from an unknown probability distribution . The nonparametric bootstrap substitutes the empirical distribution , which assigns probability to each observed record, for . Repeated values receive the combined probability of their records. A bootstrap sample is then drawn independently from , conditional on the original dataset. (blogs.helsinki.fi)
For a parameter , let the observed estimate be . A resampled dataset gives another empirical distribution, , and a bootstrap estimate . The central approximation is that the conditional distribution of resembles the sampling distribution of . Centering matters: the resamples are generated from the observed distribution, not directly from the unknown population. (blogs.helsinki.fi)
The method therefore replaces an often difficult analytical calculation with repeated simulation. Its accuracy depends on how well the fitted or empirical distribution reproduces the features relevant to the statistic. It is not an exact reconstruction of the population. (arxiv.org)
Resampling procedure
An ordinary bootstrap algorithm proceeds as follows:
- Calculate the statistic from the original observations.
- Draw indices independently and uniformly from , allowing repeated indices.
- Calculate the same statistic on the selected records.
- Repeat the resampling and calculation times, obtaining . (docs.scipy.org)
For example, from the illustrative dataset , one possible resample is . It repeats one observation and omits another. Repetition is essential: drawing all records without replacement merely rearranges the dataset and leaves order-invariant statistics unchanged.
The bootstrap standard error is the standard deviation of the replicate estimates:
The estimated bias is . These calculations concern variability and systematic displacement of the estimator, rather than the spread of individual observations. (doi.org)
Confidence intervals
Several interval constructions use the same bootstrap replicates but interpret them differently. The percentile interval takes the empirical and quantiles of the replicate estimates. For a nominal 95% interval, these are the 2.5th and 97.5th percentiles. The basic interval reflects those quantiles around the observed estimate:
The bias-corrected and accelerated, or BCa, interval adjusts the percentile levels for bias and changes in standard error as the parameter varies. These methods can produce different endpoints from identical resamples. None automatically guarantees its nominal coverage for every statistic or dataset. (docs.scipy.org)
Bootstrap methods also support hypothesis testing. Here the simulation scheme must represent the relevant null hypothesis; an unrestricted bootstrap distribution is not automatically the appropriate null distribution. (arxiv.org)
Variants and dependence
The parametric bootstrap simulates new datasets from a fitted probability model instead of selecting observed records. Each simulated dataset is processed using the same estimation procedure. This can reproduce outcomes absent from the original sample, but its validity depends on the model specification. (stat.cmu.edu)
The resampling unit must reflect the data structure. For paired measurements, shared indices preserve the pairing; independently resampling the two components destroys their association. Stratified resampling draws separately within designated strata. Such arrangements change the resampling scheme to match the sampling design. (docs.scipy.org)
For dependent time series, ordinary observation-level resampling destroys temporal relationships. A block bootstrap instead resamples contiguous stretches of observations, preserving dependence within blocks. Fixed-length and random-length block schemes are available. Their behavior depends on block length and assumptions about the underlying process, including stationarity in standard formulations. (stat.cmu.edu)
Machine-learning applications
In machine learning, bootstrap samples can serve as alternative training datasets. Bootstrap aggregating, or bagging, fits a predictor to each resample and combines predictions through averaging or voting. It is an ensemble method, with bootstrap sampling supplying the variation among training sets rather than directly constructing an uncertainty interval. (stat.berkeley.edu)
Bagging is particularly useful for unstable learning procedures whose predictions change substantially under small changes in their training data. Decision-tree learning is a principal example examined in Leo Breiman’s original work. Repeated observations remain complete records: their predictors and response values are resampled together. (doi.org)
Accuracy and computational limitations
Bootstrap validity is problem-specific. Under suitable regularity conditions, it can approximate distributions and interval coverage more accurately than first-order asymptotic methods; that improvement is not universal. Some nonregular estimators have inconsistent ordinary bootstrap distributions, as demonstrated for the Grenander estimator of a decreasing density. (arxiv.org)
Two distinct errors remain: finite- simulation error and the error of substituting an estimated distribution for the population distribution. Increasing reduces the former, generally with diminishing returns, but does not eliminate the latter or compensate for inappropriate modeling. Bootstrap replicates are simulated datasets, not additional independent population observations. (stat.cmu.edu)