Deep learning is a branch of machine learning concerned with models that learn representations through multiple successive processing layers. Most implementations use artificial neural networks, whose adjustable parameters are learned from data. Deep learning belongs to the broader field of artificial intelligence and is distinguished by its emphasis on learning intermediate representations rather than relying entirely on manually specified features. “Deep” refers to the composition of transformations or concepts, not to human-like understanding; there is no universally agreed minimum depth that makes a model deep. (deeplearningbook.org)
Representation and computation
The central principle is representation learning: a system learns how to express its inputs in forms useful for a task. In image recognition, early layers may respond to edges, intermediate layers to combinations of shapes, and later layers to structures useful for distinguishing objects. This hierarchy is an explanatory example, rather than a guarantee that every layer corresponds to a recognizable human concept. Learned representations can reduce dependence on manual feature engineering. (deeplearningbook.org)
A basic feedforward network composes transformations of the form
[ h^{(\ell)}=\sigma!\left(W^{(\ell)}h^{(\ell-1)}+b^{(\ell)}\right), ]
where (h^{(\ell)}) is the representation at layer (\ell), (W^{(\ell)}) contains weights, (b^{(\ell)}) contains biases, and (\sigma) is an activation function. Nonlinear activations allow the network to represent nonlinear relationships. Without intervening nonlinear operations, successive affine transformations collapse into one affine transformation, eliminating much of the representational advantage of depth. (deeplearningbook.org)
Depth and width describe different properties: depth concerns successive transformations, whereas width concerns the size of layers. Their usefulness depends on the function being learned, the architecture, and the available data; adding layers alone does not ensure better performance. (deeplearningbook.org)
Training and learning settings
Training adjusts parameters using training data and a loss function that expresses the learning objective. In supervised learning, examples contain target labels or values. Classification commonly uses cross-entropy, while numerical prediction may use squared error. A forward computation produces predictions and a loss; backpropagation then efficiently computes derivatives of that loss with respect to parameters. Backpropagation computes gradients—it is not itself the parameter-update rule. (deeplearningbook.org)
An optimizer such as stochastic gradient descent updates parameters, usually using small batches of examples. The learning rate controls update size. Training deep networks generally involves nonconvex optimization, so results depend on initialization, optimization settings, and the conditioning of the problem. Gradients can become too small or too large across long computational paths, making learning difficult. (deeplearningbook.org)
Deep learning is not restricted to labeled examples. Self-supervised learning constructs learning targets from the data themselves, such as predicting masked words in text. A pretrained model can subsequently undergo fine-tuning on a specific task. BERT demonstrated this approach by learning bidirectional representations from unlabeled text and adapting them to language-understanding benchmarks. (arxiv.org)
Deep networks can also be combined with reinforcement learning, where an agent learns from rewards obtained through interaction. In this setting, a network may approximate a value function or another component of the decision process rather than simply classify independent examples. (nature.com)
Major architectures
Convolutional neural networks use local connectivity and shared parameters to process structured inputs such as images. Their design exploits repeated spatial patterns while reducing the number of independently learned weights. Recurrent neural networks process sequences using a state carried between steps; gated variants help preserve information over longer sequences. Both families have been used extensively in visual and speech processing. (nature.com)
The Transformer architecture, introduced in 2017, uses attention-based processing rather than the recurrence of conventional sequence models. Its self-attention operations allow positions in a sequence to incorporate information from other positions. The original Transformer was evaluated on machine translation and supported more parallelizable training than the recurrent architectures it compared against. (arxiv.org)
Architectures encode different assumptions about inputs and their relationships. Consequently, deep learning denotes a family of modeling methods, not a single network design or training algorithm. (deeplearningbook.org)
Historical development and applications
Deep learning developed from earlier neural-network research and successive advances in representation learning, optimization, data availability, and computing hardware. Backpropagation became prominent in neural-network research during the 1980s, while research on deep models gained renewed momentum in the 2000s. Its development was cumulative rather than the result of one invention. (deeplearningbook.org)
A notable milestone was AlexNet in 2012. Its deep convolutional network achieved substantially improved image-classification results and used graphics processing units to accelerate computation. Its performance helped establish deep networks as an important approach in computer vision. The original paper also demonstrated the practical importance of nonlinear activations and methods for reducing overfitting. (proceedings.neurips.cc)
Applications include object recognition, speech recognition, and natural language processing. These tasks benefit from models that learn representations directly from complex inputs, including pixels, acoustic signals, and text. (nature.com)
Generalization and limitations
Successful training is distinct from generalization to unseen examples. Overfitting occurs when a model fits training examples without achieving correspondingly good performance on new data. Regularization methods—including parameter penalties, dropout, data augmentation, and early stopping—seek to improve generalization. Evaluation separates data used to fit parameters from data used to select model settings and estimate performance. (deeplearningbook.org)
Predictive accuracy does not automatically imply robustness or an interpretable explanation. An adversarial example is an input modified to induce an incorrect prediction; research has shown that carefully chosen, visually imperceptible image perturbations can fool trained networks. These findings distinguish ordinary test-set performance from resistance to deliberately constructed inputs and show that learned internal representations need not align neatly with human interpretations. (arxiv.org)