Supervised learning is a machine learning paradigm in which a model learns from examples pairing inputs with known target outputs. The objective is to predict outputs for previously unseen inputs, rather than merely reproduce the training examples. Inputs may be measurements, text, images, or audio; targets may be categories, numerical values, or more complex objects. “Supervised” refers to the availability of target information during training, not necessarily to continuous human oversight. (developers.google.com)
Tasks and training examples
A supervised dataset consists of input–target pairs, conventionally written . The input contains features used for prediction, while is the label or target. Training data provide the examples from which the predictive relationship is learned. For rainfall prediction, features might include temperature, humidity, and atmospheric pressure, while the target is the recorded rainfall amount. Dataset size, diversity, and quality influence how well the resulting model generalizes. (developers.google.com)
The two principal task types are classification and regression. Classification predicts membership in categories, such as whether an email is spam. Binary classification has two classes; multiclass classification has more than two; multilabel classification permits several labels for one example. Regression predicts numerical quantities, such as travel time or housing price. More elaborate tasks can predict multiple outputs or structured objects rather than a single scalar or category. (developers.google.com)
Learning objective
A model represents a predictive function , where denotes adjustable parameters. A loss function measures disagreement between predictions and targets. Many supervised methods minimize an objective of the form
The average loss is the empirical risk. The optional penalty expresses regularization, with controlling its strength. Squared error is a common regression loss; classification can use losses such as cross-entropy or hinge loss. This formulation describes many, but not all, supervised algorithms. (fairmlbook.org)
Parameter fitting may use gradient descent or stochastic gradient descent, which estimates updates from individual examples or small batches. Optimization and generalization are distinct: reducing training loss does not necessarily reduce error on new data. Overfitting occurs when a model fits the training observations too closely to perform reliably on unseen examples. Regularization can restrict effective model complexity or discourage excessively large parameter values. (fairmlbook.org)
Models and representations
Supervised learning encompasses several algorithm families. Linear regression models numerical outputs through weighted combinations of features. Logistic regression, despite its name, is commonly used for classification. Decision trees predict through sequences of feature-based decisions; random forests combine multiple trees, while gradient boosting builds an ensemble through successive additions. (scikit-learn.org)
Other families include support vector machines, which can use kernel methods to represent nonlinear relationships; the k-nearest neighbors algorithm, which bases predictions on nearby examples; and naive Bayes classifiers, which apply probabilistic models with conditional-independence assumptions. These families differ in their representations, assumptions, and computational requirements. (scikit-learn.org)
Artificial neural networks learn parameterized transformations and can accommodate complex inputs and outputs. Their training commonly uses backpropagation to compute gradients. Input representation also matters: feature engineering includes constructing or transforming predictors, while preprocessing may encode categories, scale numerical values, or handle missing observations. These transformations form part of the learning procedure, not merely an incidental preparation step. (fairmlbook.org)
Evaluation and generalization
Evaluation separates parameter fitting from performance measurement. A training set fits the model; a validation set supports choices such as model family and hyperparameters; a held-out test set assesses the selected procedure. Cross-validation repeatedly partitions data into training and validation portions, allowing comparison across several splits. Repeatedly adjusting a model to improve its test score compromises the independence of that test. (scikit-learn.org)
Evaluation metrics depend on the task. Regression metrics include mean absolute error and mean squared error. Classification metrics include accuracy, precision, recall, and F1 score. Precision measures how many predicted positives are actually positive; recall measures how many actual positives are detected. Accuracy alone can obscure poor detection of an uncommon class, and probability estimates can require evaluation beyond the final class decisions. (scikit-learn.org)
Data leakage occurs when model construction uses information unavailable at prediction time. For example, fitting feature selection or normalization on the entire dataset before evaluation allows test observations to influence training. Grouped observations and time-dependent records also require evaluation partitions appropriate to their dependence structure: ordinary random splitting may fail to represent prediction for new groups or future periods. (scikit-learn.org)
Related paradigms and limitations
Unsupervised learning seeks structure without supplied prediction targets. Semi-supervised learning combines labeled and unlabeled examples. Self-supervised learning creates supervisory targets from the data themselves, such as reconstructing withheld portions of an input. Reinforcement learning instead concerns actions and reward signals obtained through interaction. These distinctions describe learning signals rather than mutually exclusive model architectures. (developers.google.com)
Supervised performance remains tied to the meaning and coverage of the labels. A model trained on one distribution may perform differently when inputs change, and concept drift can alter the relationship between features and targets. A strong held-out score therefore characterizes performance under the evaluation conditions; it does not establish reliability for every population, environment, or future period. (developers.google.com)