An artificial neural network (ANN) is a computational model composed of interconnected processing units whose numerical parameters determine how inputs are transformed into outputs. Widely used in machine learning, these networks draw loose inspiration from biological neurons but are not generally detailed simulations of nervous systems. They can learn intermediate representations rather than depend entirely on manually specified features. Networks with multiple stages of learned processing are central to deep learning. (deeplearningbook.org)
Historical development
In 1943, Warren McCulloch and Walter Pitts described networks of idealized, all-or-none units and showed how their activity could express logical operations. Their model established a mathematical connection between neural organization and computation; it was not a modern data-trained network. The perceptron, developed by Frank Rosenblatt in the 1950s, subsequently introduced an influential approach to learning connection weights for classification. (doi.org)
A 1986 paper by David Rumelhart, Geoffrey Hinton, and Ronald Williams demonstrated how error propagation could train hidden units to represent useful features. It helped establish multilayer learning as a practical research direction, although gradient-based error propagation had earlier antecedents. In 2012, AlexNet demonstrated the effectiveness of a large convolutional network for ImageNet classification, combining GPU computation, nonlinear units, and dropout. The 2017 Transformer paper introduced an attention-based architecture without recurrent or convolutional layers for sequence transduction. (nature.com)
Units, layers, and representations
A common artificial unit computes
where are inputs, are connection weights, is a bias parameter, and is an activation function. Nonlinear activations allow successive layers to express relationships that a purely affine model cannot. Without intervening nonlinear operations, a chain of affine transformations remains an affine transformation. (deeplearningbook.org)
In a feedforward network, information passes from an input layer through hidden layers to an output layer without directed cycles. A multilayer perceptron typically uses densely connected layers. Depth describes successive processing stages, while width describes the number of units within a layer. Learned hidden activations provide representations that later layers combine, forming a basis for representation learning. (deeplearningbook.org)
Universal approximation results establish that certain networks can approximate continuous functions on compact domains arbitrarily closely, given sufficient capacity. These are statements about representational possibility, not guarantees that training will discover the required parameters or that predictions will be accurate outside the observed data. (deeplearningbook.org)
Learning and inference
Training adjusts parameters using training data and an objective. In supervised learning, examples include target outputs, such as class labels or numerical measurements. Self-supervised learning derives learning targets from the data itself, while unsupervised learning can seek structure without externally supplied labels. Neural networks are therefore a model family, not a single learning procedure. (deeplearningbook.org)
A loss function measures how outputs differ from the training objective. Backpropagation efficiently computes derivatives of that loss with respect to parameters by applying the chain rule backward through the computation. It calculates gradients; an optimizer performs the parameter updates. Stochastic gradient descent and related methods commonly estimate updates from small batches rather than the entire dataset. Learning rate, initialization, and numerical conditioning affect training behavior. (nature.com)
After training, neural network inference applies the learned computation to inputs. In ordinary deployment, producing an output does not itself update the weights. This separates the potentially expensive parameter-learning process from repeated prediction, although some systems incorporate additional adaptation mechanisms. (deeplearningbook.org)
Major architectures
Different architectures encode different assumptions about the organization of data:
- Convolutional neural networks apply shared filters across spatial or temporal grids. Local connectivity and parameter sharing make them useful for images and other structured signals. (deeplearningbook.org)
- Recurrent neural networks maintain states across sequence positions, sharing parameters over successive steps. Long short-term memory networks use gates and memory states to help model dependencies over longer sequences. (deeplearningbook.org)
- Transformers use an attention mechanism to combine information from different sequence positions. The original architecture paired attention with position-wise feedforward layers and positional information, enabling substantial parallel computation during training. (arxiv.org)
- Autoencoders learn an encoding and a reconstruction of their inputs. Constraints on their representations or training objectives encourage them to capture useful structure rather than simply copy each input unchanged. (deeplearningbook.org)
These families are not mutually exclusive: a larger system can combine convolutional, recurrent, attention-based, and densely connected components. (deeplearningbook.org)
Generalization and limitations
The central evaluation problem is generalization: performance on previously unseen examples. A network may exhibit overfitting, achieving low training error while capturing patterns that do not transfer. Regularization methods include parameter penalties, data augmentation, early stopping, and dropout, which randomly suppresses selected activations during training. A validation set supports model selection, while a separate test set estimates performance after those choices. (deeplearningbook.org)
Large networks can require substantial computation, memory, and data; graphics processing units accelerate many of their numerical operations. Predictive success remains task-dependent, and low training loss does not establish reliable behavior on unfamiliar inputs. Networks can also be vulnerable to adversarial examples—deliberately constructed perturbations that cause erroneous predictions despite appearing small or inconspicuous to a human observer. Such behavior demonstrates that learned decision boundaries need not match human perceptual judgments. (proceedings.neurips.cc)