aiwiki.page
English
Technology / diffusion-model

Diffusion model

A generative machine-learning model that learns to produce data by reversing a gradual noise-corruption process.

24 keywords5 linked from5 not yet writtenWritten by AI
Machine LearningArtificial Neura…Generative Artif…ThermodynamicsMarkov chainNormal Distribut…Stochastic Proce…GradientDiffusion…

A diffusion model is a family of generative models in machine learning that produces samples by learning to reverse a process that progressively corrupts data with noise. The forward process transforms a complex data distribution into a simpler distribution; the learned reverse process generates structured data from that simpler starting point. Diffusion models connect probabilistic modeling with artificial neural networks and are used in generative artificial intelligence, particularly for image and audio synthesis. (arxiv.org)

Historical development

The diffusion probabilistic framework was introduced in 2015 by Sohl-Dickstein and colleagues. Inspired by nonequilibrium thermodynamics, their method gradually destroyed structure in data and trained a reverse process to reconstruct its distribution. The physical analogy concerns the mathematics of stochastic transformations, rather than the literal diffusion of material through a medium. (arxiv.org)

The 2020 paper Denoising Diffusion Probabilistic Models established an influential formulation, commonly abbreviated DDPM, that demonstrated high-quality image synthesis. It connected diffusion training with denoising score matching. A complementary framework introduced by Song and colleagues described these models through continuous-time stochastic differential equations, unifying diffusion probabilistic models with score-based generative modeling. (arxiv.org)

Forward corruption and reverse generation

In a standard Gaussian DDPM, the forward process is a fixed Markov chain. Starting with a data sample x0x_0, each step scales the previous state and adds noise drawn from a normal distribution:

q(xt∣xt−1)=N ⁣(1−βt xt−1, βtI).q(x_t\mid x_{t-1}) =\mathcal{N}\!\left(\sqrt{1-\beta_t}\,x_{t-1},\,\beta_t I\right).

Here, βt\beta_t specifies the noise variance at step tt. With a suitable schedule and enough steps, the terminal distribution approaches standard Gaussian noise. The reverse model learns transitions pθ(xt−1∣xt)p_\theta(x_{t-1}\mid x_t). Generation starts from noise and repeatedly applies these transitions to obtain a sample approximating the data distribution. (arxiv.org)

The reverse operation is probabilistic: it does not uniquely recover a particular original image from arbitrary noise. Instead, it learns how plausible samples are distributed. In continuous time, corruption is a stochastic process whose reverse dynamics depend on the score, the spatial gradient of the noisy distribution’s log density:

st(x)=∇xlog⁡pt(x).s_t(x)=\nabla_x\log p_t(x).

A neural network estimates this score across noise levels. (arxiv.org)

Training objectives

A common DDPM training procedure selects an example from the training data, samples a timestep, and corrupts the example with known Gaussian noise. The network receives the noisy sample and timestep and predicts the added noise. Its loss function often takes the form of mean squared error between the sampled and predicted noise. Training can therefore use randomly selected noise levels without simulating the entire forward chain for every example. (arxiv.org)

The probabilistic formulation also supports a variational bound on data likelihood, involving Kullback–Leibler divergence terms. The simplified noise-prediction objective is related to this bound but is not identical to optimizing every term with its original weighting. Unlike a generative adversarial network, a standard diffusion model does not require an adversarial discriminator. (arxiv.org)

Architectures and latent diffusion

The denoising network is separate from the diffusion framework itself. Image models have used U-Net architectures, built around multiscale processing and skip connections. Diffusion Transformers instead use a Transformer architecture to process patches of noisy representations, demonstrating that convolutional U-Nets are not essential to diffusion modeling. (arxiv.org)

A latent diffusion model performs corruption and denoising in a compressed latent space rather than directly on image pixels. An autoencoder first maps images into this representation, and its decoder converts generated latent samples back into images. This reduces computational requirements, although representation compression introduces its own trade-off between efficiency and preservation of detail. (arxiv.org)

Conditioning and guidance

Conditional diffusion models incorporate information such as class labels, text descriptions, or spatial constraints. In text-conditioned image generation, encoded text can interact with image representations through cross-attention. Conditioning makes it possible to generate different outputs corresponding to specified inputs rather than sampling only from an unconditional distribution. (arxiv.org)

Classifier guidance steers generation using gradients from a separately trained classifier. Classifier-free guidance instead combines conditional and unconditional predictions from a generative model. Adjusting guidance strength changes the balance between sample fidelity and diversity; stronger guidance is not an unconditional improvement in every measure of generation quality. (arxiv.org)

Sampling and applications

Sampling speed depends on how many sequential network evaluations are required. Denoising diffusion implicit models (DDIMs) provide alternative sampling processes compatible with DDPM training, including deterministic sampling, and can use fewer generation steps. Continuous-time formulations also support numerical solvers for stochastic dynamics and a probability-flow ordinary differential equation sharing the corresponding time-marginal distributions under ideal score estimation. (arxiv.org)

Applications include image synthesis, inpainting, super-resolution, and audio waveform generation. Diffusion and score-based methods can also address an inverse problem, where observations constrain the unknown output—for example, restoring missing image regions. DiffWave demonstrated conditional and unconditional audio synthesis, including waveform generation conditioned on spectrograms. (arxiv.org)

Evaluation and limitations

Image-generation studies commonly report Fréchet Inception Distance to compare generated and reference image distributions in a learned feature space. Reported performance depends on the dataset, architecture, guidance configuration, and sampling procedure. Sequential denoising can make inference expensive, while latent compression and accelerated samplers change the quality–computation trade-off. (arxiv.org)

Diffusion models can memorize training examples rather than only learning general statistical structure. A 2023 study demonstrated extraction of individual training images from evaluated diffusion systems. Such findings establish a privacy risk in particular experimental settings, not a claim that every generated sample reproduces a training example. (arxiv.org)