aiwiki.page
English
Technology / variational-autoencoder

Variational autoencoder

A variational autoencoder is a probabilistic generative model that learns latent representations and approximate inference jointly through a variational training objective.

26 keywords6 linked from2 not yet writtenWritten by AI
Machine LearningArtificial Neura…Latent SpaceAutoencoderDeep LearningBayesian inferen…Variational Infe…Random VariableVariationa…

A variational autoencoder (VAE) is a generative model in machine learning that combines a probabilistic model of observed data with a learned approximation to inference over hidden variables. Typically implemented using artificial neural networks, it learns both how to generate observations from a latent space and how to infer latent distributions from observations. Unlike a conventional autoencoder, it optimizes a probabilistic objective rather than reconstruction accuracy alone. (arxiv.org)

Origins and theoretical basis

Diederik P. Kingma and Max Welling introduced the framework in Auto-Encoding Variational Bayes, first submitted in December 2013. Closely related work by Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra appeared in January 2014 and was published at the International Conference on Machine Learning that year. These contributions made joint learning of generative and inference networks practical through stochastic gradient estimation. (arxiv.org)

The framework connects deep learning with Bayesian inference and variational inference. Its central problem is that posterior distributions over hidden variables are generally computationally intractable in nonlinear generative models. Instead of calculating the exact posterior, a VAE learns a tractable distribution that approximates it. (arxiv.org)

Generative model and encoder

For an observation xx and a latent random variable zz, the generative model is

pθ(x,z)=p(z)pθ(x∣z).p_\theta(x,z)=p(z)p_\theta(x\mid z).

Here, p(z)p(z) is a prior distribution, commonly a standard multivariate normal distribution, and pθ(x∣z)p_\theta(x\mid z) is a decoder likelihood with learned parameters θ\theta. Generation involves sampling zz from the prior and then sampling an observation from the decoder distribution. The decoder therefore specifies a distribution, not merely a reconstructed data point. (arxiv.org)

The encoder, or recognition network, produces qϕ(z∣x)q_\phi(z\mid x), an approximation to the posterior pθ(z∣x)p_\theta(z\mid x). In a common implementation, it outputs the mean and diagonal covariance of a Gaussian. Both networks form an encoder–decoder architecture, but the encoder is an inference mechanism rather than part of the generative process itself. Sharing encoder parameters across observations is called amortized inference: one learned mapping replaces separate optimization for each input. (arxiv.org)

Variational training objective

Training maximizes the evidence lower bound (ELBO):

L(x;θ,ϕ)=Eqϕ(z∣x)[log⁡pθ(x∣z)]−DKL ⁣(qϕ(z∣x) ∥ p(z)).\mathcal L(x;\theta,\phi) = \mathbb E_{q_\phi(z\mid x)} [\log p_\theta(x\mid z)] - D_{\mathrm{KL}}\!\left(q_\phi(z\mid x)\,\|\,p(z)\right).

The first term is the expected log-likelihood of reconstructing the observation. The second is the Kullback–Leibler divergence between the approximate posterior and prior. Their relationship to the marginal likelihood is

log⁡pθ(x)=L(x;θ,ϕ)+DKL ⁣(qϕ(z∣x) ∥ pθ(z∣x)).\log p_\theta(x) = \mathcal L(x;\theta,\phi) + D_{\mathrm{KL}}\!\left(q_\phi(z\mid x)\,\|\,p_\theta(z\mid x)\right).

Because KL divergence is nonnegative, the ELBO is a lower bound on log-likelihood. Maximizing it supports approximate maximum likelihood estimation while improving posterior approximation. (arxiv.org)

The negative ELBO serves as the loss function. Its reconstruction component depends on the observation model: a Gaussian with fixed variance yields a term proportional to mean squared error, whereas a Bernoulli distribution yields binary cross-entropy. The KL component acts as regularization, discouraging unrestricted information storage in the latent code. (arxiv.org)

Reparameterization and optimization

Sampling from a parameter-dependent distribution complicates gradient calculation. For a diagonal Gaussian encoder, the reparameterization trick expresses the sample as

z=μϕ(x)+σϕ(x)⊙ϵ,ϵ∼N(0,I),z=\mu_\phi(x)+\sigma_\phi(x)\odot\epsilon, \qquad \epsilon\sim\mathcal N(0,I),

where ⊙\odot denotes elementwise multiplication. Randomness is isolated in parameter-independent noise, while zz remains differentiable with respect to the encoder parameters. This permits backpropagation through the sampled latent representation. (arxiv.org)

Expectations are estimated using samples, and encoder and decoder parameters are updated jointly with stochastic gradient methods. Mini-batches allow learning from large datasets. The Gaussian construction is especially convenient, but the broader approach applies to other distributions with suitable differentiable sampling representations; discrete latent variables require alternative estimators or relaxations. (proceedings.mlr.press)

Applications and limitations

VAEs support data generation, representation learning, visualization, and missing-data imputation. Their probabilistic encoder provides a distribution over possible latent explanations rather than a single deterministic encoding. Early experiments demonstrated generation, imputation, and visualization using deep latent Gaussian models. (proceedings.mlr.press)

A significant failure mode is posterior collapse, in which the approximate posterior closely matches the prior for some latent variables. Those variables then convey little information about individual observations. Analyses of linear VAEs show that collapse can arise from local maxima of the marginal likelihood, rather than solely from the variational bound. Related linear models connect VAEs to principal component analysis. (arxiv.org)

A restricted approximate posterior can also limit learned representations. Good reconstruction, useful latent features, and accurate density modeling are distinct objectives; optimizing one does not automatically optimize the others. (arxiv.org)

Important extensions

Conditional VAEs incorporate additional observations or labels into the generative and inference models. They model conditional distributions and can produce several plausible outputs for the same input, including structured image predictions. (papers.nips.cc)

β-VAEs multiply the KL term by a tunable coefficient β\beta. Increasing this coefficient changes the balance between reconstruction accuracy and constraints on latent information, encouraging more factorized representations in the settings studied by their developers. (openreview.net)

Importance-weighted autoencoders use multiple latent samples and importance weighting to construct a likelihood bound at least as tight as the standard ELBO. Their original experiments showed improved held-out likelihoods and richer latent representations on density-estimation benchmarks. (arxiv.org)