A convolutional neural network (CNN) is a type of artificial neural network designed to process data arranged on regular grids, including images and sampled signals. Its defining operation applies learned filters across different positions, allowing the same pattern detector to operate throughout an input. CNNs are important architectures in deep learning and computer vision, combining local connectivity, shared parameters, and multiple processing layers to construct increasingly complex representations. Unlike systems built around fixed, manually designed image descriptors, they can learn useful features directly from examples. (deeplearningbook.org)
Convolution and parameter sharing
An image is typically represented as a tensor with height, width, and channel dimensions; a color image commonly has three channels. A convolutional layer contains trainable filters, also called kernels. Each filter combines values from a local input region, usually across all input channels, to produce one value in an output feature map. Moving the filter across the input produces a spatial map of responses. Different filters generate different output channels. (docs.pytorch.org)
For a two-dimensional layer with unit stride, the operation can be written as
where is the input, contains filter weights, is an output-channel bias, and is the output. Despite the conventional name, implementations generally compute cross-correlation: they do not reverse the kernel as mathematical convolution does. For learned filters, this naming distinction does not change the class of mappings the layer can represent. (docs.pytorch.org)
Parameter sharing means that a filter uses identical weights at every spatial position. Consequently, the number of learned weights need not grow with image area. Local connectivity restricts each output to a neighborhood rather than connecting it to every input value. These constraints encode assumptions about spatial structure and make convolutional layers more economical than comparable fully connected layers. (deeplearningbook.org)
Network structure
A typical CNN interleaves convolutional layers with nonlinear activation functions, such as the rectified linear unit, . Nonlinearity allows successive layers to represent transformations beyond a single linear mapping. Early image-processing layers may respond to edges or color contrasts, while deeper layers combine these responses into more complex patterns. This hierarchical representation learning reduces reliance on manual feature engineering. (leon.bottou.org)
Several settings control spatial processing:
- Stride specifies how far a filter moves between applications; larger strides commonly reduce output resolution.
- Padding extends the input boundary, often with zeros, to control output size.
- Dilation spaces the kernel’s sampled positions farther apart, increasing its spatial reach without adding kernel weights. (en.d2l.ai)
Pooling aggregates nearby responses, usually by taking their maximum or average. Standard pooling has no learned weights and normally operates independently on each channel. With downsampling, it reduces spatial resolution and computational cost, but also discards detailed positional information. Pooling is common rather than essential: other architectures reduce resolution through strided convolution. (github.com)
For image classification, the final feature maps may feed fully connected layers or be spatially averaged before a classifier. A softmax function can convert the resulting scores into a distribution over classes. Other CNNs retain spatial outputs instead of producing a single image-level prediction. (deeplearningbook.org)
Training and reuse
CNNs are frequently trained through supervised learning, using labeled training data and a loss function that measures disagreement between predictions and targets. Backpropagation computes derivatives of the loss with respect to filter weights and other parameters. An optimizer, commonly based on stochastic gradient descent, then updates those parameters using batches of examples. Feature extraction and prediction can therefore be optimized together, end to end. (leon.bottou.org)
Large networks can exhibit overfitting. Regularization methods include weight decay and dropout, while data augmentation creates transformed training examples through operations such as cropping and horizontal reflection. Such transformations must preserve the target meaning. Batch normalization introduces normalization within the network and can improve optimization behavior. (papers.nips.cc)
In transfer learning, representations learned for one task are reused for another. A pretrained classification network can, for example, provide the starting representation for pixel-level prediction, with fine-tuning adapting its parameters to the new objective. (arxiv.org)
Historical development
CNN development drew on hierarchical models of visual processing and on trainable systems for pattern recognition. During the late 1980s and 1990s, convolutional networks trained with gradient-based methods became practical for handwritten-character recognition. LeNet-5, described in a 1998 paper by Yann LeCun and colleagues, combined convolution, subsampling, and classification layers for document-recognition tasks. (leon.bottou.org)
In 2012, AlexNet demonstrated strong large-scale image-classification results on the ImageNet competition. Its winning entry achieved a 15.3% top-five test error, compared with 26.2% for the runner-up. Training used graphics processing units, rectified activations, and dropout. Later, residual networks introduced shortcut connections that add a block’s input to its learned transformation, enabling substantially deeper CNNs to be optimized successfully. The residual-network paper presented at CVPR 2016 included a 152-layer ImageNet model. (papers.nips.cc)
Applications and limitations
Beyond classification, CNNs support object localization and detection. Fully convolutional architectures perform semantic segmentation, assigning categories to individual pixels and combining coarse semantic features with finer spatial information. Convolution also applies to one-dimensional sampled signals and other grid-structured inputs. (cv-foundation.org)
Shared filters provide translation equivariance: under suitable boundary and sampling conditions, shifting the input shifts the feature responses correspondingly. This is not the same as complete translation invariance, nor does ordinary convolution guarantee invariance to rotation or changes of scale. CNN predictions can also be vulnerable to adversarial examples—deliberately constructed, sometimes visually imperceptible input perturbations that cause misclassification. (deeplearningbook.org)