A latent space is the set of possible values taken by unobserved variables or internal representations in a statistical or computational model. In machine learning, it provides an alternative description of observed data: an image, for example, may be represented by coordinates encoding aspects of its appearance rather than by its pixels directly. “Latent” indicates that these coordinates are not directly observed as measurements in the original data. The term encompasses both probabilistic models of hidden factors and learned codes in representation learning; their coordinates need not have recognizable meanings. (deeplearningbook.org)
Mathematical formulation
For an observation , a deterministic encoder produces a representation , where denotes learned parameters. Its latent space is the domain of possible codes . Often , a vector space with coordinates, although models can also use discrete variables or structured combinations. The representation is frequently smaller than the original input, but reduced dimensionality is not a defining requirement: overcomplete representations may contain more coordinates than their inputs. (deeplearningbook.org)
In a probabilistic formulation, is a random variable, and the model specifies a joint probability distribution
Here is a prior distribution, while describes observations conditional on the latent state. The observed-data distribution is obtained by integrating out , or summing over its possible values when it is discrete. Inferring from an observation instead concerns the posterior distribution . Thus, inference and generation traverse the relationship between observations and latent variables in opposite directions. (arxiv.org)
Learning latent representations
Latent representations can be constructed by linear methods or learned through nonlinear models. Principal component analysis represents centered observations through coordinates along selected directions of greatest variance. Its reduced representation lies in a linear subspace. Nonlinear encoders can instead capture more complicated relationships between observations and their representations. (deeplearningbook.org)
An autoencoder uses an encoder–decoder architecture: the encoder maps to a code , and the decoder produces a reconstruction . Training minimizes a reconstruction loss function, potentially together with additional constraints. A narrow bottleneck encourages dimensionality reduction, while sparsity, noise-based training, and other forms of regularization can encourage useful representations even without a narrow bottleneck. Reconstruction quality alone does not establish that the code captures meaningful explanatory factors. (deeplearningbook.org)
A variational autoencoder gives the code an explicitly probabilistic interpretation. Its encoder typically specifies an approximate posterior , rather than only one point. Training commonly maximizes an evidence lower bound:
The first term rewards explaining the observation; the second, a Kullback–Leibler divergence, penalizes departure from the prior. A standard normal distribution is a common prior choice, but not a necessary one. (arxiv.org)
Geometry and interpolation
A latent space has a coordinate system, but coordinates alone do not determine a useful notion of similarity. Euclidean distance treats all coordinate directions uniformly; a nonlinear decoder can nevertheless transform equal latent displacements into very different changes in the output. Consequently, geometric proximity and semantic similarity should not be treated as interchangeable properties. (proceedings.mlr.press)
For a differentiable decoder , its Jacobian matrix can induce the local metric
When appropriate regularity and rank conditions hold, this describes lengths inherited from the output space and connects latent representations with differential geometry. A geodesic under this metric need not be a straight line in latent coordinates. (openaccess.thecvf.com)
Interpolation evaluates intermediate codes between two endpoints. Linear interpolation uses ; geometry-aware methods can follow curved paths instead. Smooth-looking transitions can reveal structure learned by a generator, but neither smoothness nor plausible endpoints guarantees that intermediate outputs follow the learned data manifold or preserve a particular attribute. (proceedings.mlr.press)
Generative models and applications
In a generative adversarial network, a generator transforms a sampled latent input into a synthetic observation. The original framework learns this mapping through competition with a discriminator and does not require an encoder that maps observations back to latent codes. Latent space therefore need not be created through compression or reconstruction. (arxiv.org)
Latent-space computation also supports generative artificial intelligence through latent diffusion models. These systems learn an autoencoder representation, train a diffusion model in that representation, and decode generated latents into images. Moving the denoising process away from full-resolution pixels reduces computational demands, with the representation balancing compression against preservation of visual detail. Conditioning mechanisms can incorporate text or other inputs. (openaccess.thecvf.com)
Beyond generation, learned representations can support prediction and transfer across tasks or modalities. Their usefulness depends on which information they retain and which distinctions they make accessible to downstream models; a representation effective for reconstruction is not automatically optimal for every predictive task. (deeplearningbook.org)
Interpretability and identifiability
A disentangled representation seeks to separate underlying factors of variation into distinct coordinates or coordinate groups. This is stronger than simply obtaining a compact code. Under broad latent-variable assumptions, different transformations of the hidden variables can explain the same observed distribution, so observations alone may not uniquely identify the intended factors. Research on unsupervised learning has demonstrated that disentanglement cannot generally be recovered without additional assumptions about the model and data. Coordinate labels such as “pose” or “style” therefore require evidence rather than being intrinsic properties of latent space. (proceedings.mlr.press)