Image processing is the application of operations to images to modify their appearance, recover information, or prepare them for analysis, storage, and transmission. Digital image processing uses computational methods to manipulate numerical image representations. It overlaps with signal processing and computer vision: image processing commonly transforms image data, while computer vision seeks to interpret scenes, objects, and activities. Their boundaries are not absolute, and practical systems often combine enhancement, geometric correction, measurement, and recognition. (mathworks.com)
Image representation
A digital image consists of spatially arranged samples called pixels. A grayscale image can be represented as a matrix whose entries encode intensity; a color image typically includes several channels. Images with additional spatial, spectral, or temporal dimensions can be represented as multidimensional arrays or tensors. The interpretation of their values depends on the acquisition system, numerical format, and associated metadata. (scikit-image.org)
Bit depth determines how many discrete values a sample can encode. An unsigned eight-bit channel provides 256 possible values, whereas a sixteen-bit channel provides 65,536. Floating-point representations accommodate intermediate computations and wider numerical ranges, although software libraries may impose particular conventions. Converting between formats can change scaling or discard precision. Spatial dimensions and intensity precision describe different aspects of an image and should not be confused. (scikit-image.org)
A color space specifies how numerical components represent color. RGB uses red, green, and blue channels; HSV expresses hue, saturation, and value. Color-space conversion can make particular operations easier, such as adjusting brightness separately from hue. Channel ordering also matters: some software uses RGB, while other software commonly stores the same components in BGR order. (github.com)
Enhancement and filtering
Enhancement modifies images to make features more visible or more suitable for subsequent analysis. Contrast adjustment remaps intensity values; histogram equalization redistributes them according to their frequency. Enhancement is purpose-dependent: an operation that improves visual presentation does not necessarily preserve the original numerical measurements. Aggressive intensity stretching can also make background noise appear significant. (mathworks.com)
Spatial filters compute output values from neighborhoods of input pixels. Linear filtering often uses convolution, combining nearby values with a weighted kernel. Gaussian smoothing suppresses fine-scale variation, while sharpening emphasizes changes in intensity. Nonlinear methods include median and bilateral filters; bilateral filtering incorporates both spatial proximity and intensity similarity to reduce noise while retaining edges. Boundary handling is necessary because neighborhoods near an image’s borders extend beyond available samples. (docs.opencv.org)
Frequency-domain processing uses transformations such as the Fourier transform to express images in terms of spatial frequencies. Filtering in this representation can suppress periodic interference or emphasize particular scales. Mathematical morphology instead probes image shapes with a structuring element. Erosion and dilation, and their combinations as opening and closing, support operations such as removing small structures and modifying region boundaries. (mathworks.com)
Geometry, segmentation, and measurement
Geometric operations change the spatial arrangement of image content through scaling, rotation, translation, or perspective transformation. Output sample locations generally do not coincide exactly with input locations, making interpolation necessary. Nearest-neighbor, bilinear, and bicubic interpolation differ in computational cost and the way they estimate values between samples. Downsampling may require antialiasing to prevent fine patterns from becoming misleading coarse patterns. (docs.opencv.org)
Image registration aligns images within a common coordinate system, using corresponding features, control points, or intensity relationships. Image segmentation partitions an image into regions through methods such as thresholding and region-based analysis. Semantic segmentation additionally assigns category labels to pixels. After segmentation, feature extraction can produce measurements of region size, shape, intensity, or texture, converting image content into quantitative descriptors. (mathworks.com)
Restoration and learned methods
Restoration estimates an underlying image from observations degraded by blur, noise, or other acquisition effects. Unlike enhancement, it usually relies on a model connecting the original scene to the recorded data. Deblurring is an inverse problem: different candidate originals may explain similar observations, and direct inversion can amplify noise. Regularization introduces additional constraints or preferences, often within a mathematical optimization formulation. (mathworks.com)
Machine learning provides another approach, learning transformations from examples rather than specifying every operation manually. Deep learning systems, including convolutional neural networks, perform denoising, deblurring, enhancement, and super-resolution. Their behavior depends on training data and the selected loss function. For example, pixel-error objectives and perceptual objectives can favor different outputs; a visually convincing reconstruction is not necessarily an exact recovery of the original scene. (arxiv.org)
Compression, evaluation, and implementation
Image compression reduces storage and transmission requirements. Lossless compression permits exact reconstruction of encoded sample values; lossy compression accepts changes to those values in exchange for reduced data size. Standards support different combinations of these objectives: JPEG-LS provides lossless and near-lossless coding, while JPEG 2000 supports both lossy and lossless modes. (jpeg.org)
Evaluation depends on the intended task. Mean squared error measures numerical disagreement with a reference image, but does not fully capture perceived quality. Structural similarity compares local image structure alongside luminance and contrast. Implementations must also account for memory, numerical formats, and processing speed. Common software includes OpenCV, scikit-image, and MATLAB’s image-processing tools; suitable operations can execute on a graphics processing unit to accelerate demanding workflows. (cns.nyu.edu)