A random forest is a machine-learning method that combines many randomized decision trees into one predictive model. It belongs to ensemble learning and is most commonly used for supervised learning, including classification and regression. Each tree is trained with randomness introduced into its data or construction, and their predictions are aggregated. The familiar formulation combines bootstrap sampling with random feature selection at each split, reducing dependence among trees and often improving predictive stability over a single tree. (doi.org)
Development and basic principle
Leo Breiman’s 2001 paper “Random Forests” established a general framework for ensembles of tree predictors governed by independently sampled random vectors. It analyzed how prediction error depends on the strength of individual trees and the correlation between their errors. The framework built on bootstrap aggregating, or bagging, and earlier work on randomized tree construction and random feature subsets. Although bootstrap sampling is characteristic of the standard algorithm, the broader definition encompasses other ways of randomizing trees. (doi.org)
A deep decision tree can fit complex patterns but may change substantially when its training data change. This instability is associated with high variance and overfitting. Averaging trees can reduce variance, especially when their errors are not strongly correlated. Random feature selection prevents the same dominant predictors from controlling every tree, encouraging diversity rather than merely producing repeated versions of one model. (scikit-learn.org)
Training and prediction
For a standard forest trained on observations and features, tree construction follows three main steps:
- Draw a bootstrap sample of observations, usually containing draws with replacement. Some observations therefore appear repeatedly, while others are omitted.
- At each node, randomly select a subset of candidate features. Search those candidates for the split that best improves the chosen criterion.
- Continue splitting until a stopping condition is reached. Classical random forests commonly use deep, unpruned trees, although implementations permit restrictions on their size. (scikit-learn.org)
Classification trees commonly use Gini impurity or entropy to select splits. Regression trees often use a loss function based on mean squared error, with predictions formed from average target values in terminal leaves. Trees are fitted independently rather than successively correcting earlier trees. (stat-www.berkeley.edu)
For regression, a forest of trees typically predicts
where is tree ’s prediction. Breiman’s classification formulation uses majority voting. Some implementations instead average the trees’ class probability estimates and select the class with the largest average; these procedures need not produce identical results. (stat-www.berkeley.edu)
Parameters and generalization
Important hyperparameters include the number of trees, the number of candidate features per split, maximum tree depth, minimum leaf size, and bootstrap sample size. Depth and leaf-size restrictions control individual-tree complexity and act as forms of regularization. Smaller feature subsets generally increase diversity but can also weaken individual trees, illustrating a bias–variance tradeoff. (scikit-learn.org)
Increasing the number of trees reduces the randomness associated with a finite ensemble, while increasing training time, prediction time, and storage. Breiman showed that classification generalization error converges to a limit as the forest grows. This does not mean a forest cannot overfit: its performance still depends on the data, tree construction, and model-selection procedure. Adding trees does not remedy unsuitable features or an unrepresentative dataset. (doi.org)
Out-of-bag evaluation
An observation’s probability of being omitted from an -draw bootstrap sample is
which approaches , approximately 36.8%. Consequently, roughly one-third of observations are absent from each tree’s sample. These are its out-of-bag observations. For each training observation, predictions can be aggregated using only trees that did not train on it, yielding an out-of-bag error estimate without reserving a separate evaluation subset. (stat.berkeley.edu)
Out-of-bag evaluation does not automatically prevent data leakage or provide an independent final assessment after extensive tuning. Cross-validation and a separate test set serve distinct evaluation roles. When observations share subjects or other groups, or form a time series, evaluation must account for that dependence rather than treating all rows as interchangeable independent examples. (scikit-learn.org)
Feature importance and interpretation
Random forests can support feature selection through estimates of predictor importance. Impurity-based importance sums the weighted improvements attributable to each feature across tree splits and averages them across the forest. Such scores describe the fitted model’s use of features, but can favor variables with many possible split points and reflect training-set patterns that do not generalize. (scikit-learn.org)
Permutation feature importance measures the change in predictive performance after a feature’s values are shuffled. It can be computed on held-out data or through out-of-bag predictions. Correlated features complicate interpretation: shuffling one predictor may cause little deterioration because another supplies similar information. A low individual score therefore does not necessarily imply that the information represented by that predictor is unimportant. (github.com)
Computational properties and limitations
Independent tree fitting makes random forests suitable for parallel computing, although communication overhead limits speedup. Large forests of deep trees can require substantial memory and prediction work. Unlike gradient boosting, which constructs trees sequentially to improve the existing ensemble, random forests primarily combine independently randomized predictors. (scikit-learn.org)
Standard regression forests average terminal-leaf predictions. When those leaves contain means of observed targets, forest predictions remain within the training target range; this follows directly from averaging and limits extrapolation beyond observed outcomes. Random forests capture nonlinear relationships through partitions, but do not automatically learn an extrapolating trend of the kind explicitly represented by linear regression. (scikit-learn.org)