A multilayer perceptron (MLP) is an artificial neural network consisting of an input layer, one or more hidden layers, and an output layer. In its conventional form, successive layers are fully connected: every unit in one layer has a trainable connection to every unit in the next. Information flows forward without recurrent connections. MLPs are used in machine learning to learn nonlinear relationships, particularly for classification and regression, and provide a basic architecture for deep learning. (scikit-learn.org)
Architecture and computation
An MLP receives a numerical input vector and transforms it through successive layers. For a hidden layer (l), the computation can be written as
[ h^{(l)}=\phi_l!\left(W^{(l)}h^{(l-1)}+b^{(l)}\right), \qquad h^{(0)}=x. ]
Here, (W^{(l)}) is a weight matrix, (b^{(l)}) is a bias vector, and (\phi_l) is an activation function, usually applied separately to each component. The operation inside the parentheses is an affine map. Layer width refers to its number of units; depth describes the number of successive computational layers. Counting conventions differ: an architecture with an input, hidden, and output layer may be called a three-layer network, or a two-layer network when only trainable layers are counted. (d2l.ai)
Nonlinear activations are essential. Composing affine transformations without intervening nonlinearities produces another affine transformation, regardless of the number of layers. Common activations include the logistic sigmoid, hyperbolic tangent, and rectified linear unit (ReLU), defined by (\max(0,z)). Hidden layers therefore enable representation learning: intermediate features are learned rather than prescribed individually. (d2l.ai)
For adjacent layers of widths (n_{l-1}) and (n_l), a fully connected layer with biases contains (n_ln_{l-1}+n_l) trainable parameters. Consequently, widening layers can substantially increase storage and computation. This parameter count follows directly from the dimensions of the weight matrix and bias vector. (en.d2l.ai)
Outputs and training
In supervised learning, an MLP learns from input–target pairs in training data. Its output transformation depends on the task. Regression commonly uses an unrestricted affine output and mean squared error. Binary classification can use a sigmoid output, while mutually exclusive multiclass classification commonly uses a softmax function with cross-entropy loss. Output activations and losses are selected together to express the intended prediction problem. (scikit-learn.org)
Training adjusts weights and biases to reduce a loss function, often averaged across examples and supplemented by a penalty. Backpropagation computes parameter derivatives efficiently by applying the chain rule backward through the network. It is a method for calculating gradients, not the complete learning algorithm: an optimizer uses those gradients to change the parameters. (deeplearningbook.org)
Common optimizers include stochastic gradient descent and Adam. Mini-batch training estimates updates from subsets of examples. The learning rate controls update size, while layer widths, depth, batch size, and activation choices are hyperparameters. Training objectives are generally nonconvex, so initialization and optimizer settings can affect the solution reached; ordinary gradient-based training does not guarantee a global optimum. (scikit-learn.org)
Expressive power
A single perceptron implements a linear decision boundary and cannot represent every classification problem. The exclusive OR problem illustrates this limitation: its two classes cannot be separated by one straight line in the original two-dimensional input space. An MLP can transform the inputs through hidden units so that an output unit separates the resulting representations. (deeplearningbook.org)
Universal approximation results establish that, with suitable nonlinear activations and sufficiently many hidden units, a network with one hidden layer can approximate any continuous function on a compact subset of finite-dimensional Euclidean space to arbitrary accuracy. This is an existence statement about representational capacity. It does not specify an efficient training procedure, guarantee that finite data reveal the desired function, or ensure successful predictions outside the sampled region. Additional depth can represent some functions more economically than a shallow network. (deeplearningbook.org)
Generalization and numerical behavior
Large MLPs may exhibit overfitting, fitting training examples without performing comparably on unseen data. Regularization methods include weight penalties, which discourage large parameter values, and dropout, which randomly suppresses units during training. Standard dropout uses the full network at prediction time with appropriate scaling, rather than retaining the training-time random omissions. These methods constrain learning but do not guarantee improved performance on every dataset. (classic.d2l.ai)
Early stopping limits training when validation performance ceases to improve. Feature scaling also affects optimization because inputs with very different numerical ranges can make training more difficult. Architecture, regularization strength, and training duration influence the balance between fitting observed examples and generalizing beyond them. (scikit-learn.org)
Deep networks can encounter the vanishing gradient problem or excessively large gradients as derivatives propagate through successive layers. Saturating activations, weight magnitudes, and depth influence this behavior. Random initialization breaks symmetry between hidden units, while initialization schemes based on incoming and outgoing layer widths help control the scale of activations and gradients. (classic.d2l.ai)
Historical development and architectural role
The 1986 paper Learning representations by back-propagating errors, by David Rumelhart, Geoffrey Hinton, and Ronald Williams, demonstrated how error-driven weight adjustments allow hidden units to acquire useful internal features. It was an influential presentation of multilayer-network learning, rather than the first appearance of every underlying differentiation technique. (nature.com)
MLPs also serve as components within larger architectures. In the original Transformer architecture, each encoder and decoder layer includes a position-wise feedforward network comprising two affine transformations separated by a ReLU. The same transformation is applied independently at each sequence position, complementing attention, which exchanges information across positions. (arxiv.org)