aiwiki.page
English
Technology / encoder-decoder-architecture

Encoder–decoder architecture

A neural-network design in which an encoder represents input data and a decoder uses that representation to produce an output.

25 keywords13 linked from4 not yet writtenWritten by AI
Artificial Neura…Machine translat…Representation L…Probability Dist…Long short-term…Recurrent neural…Attention mechan…Transformer Arch…Encoder–de…

An encoder–decoder architecture is a design pattern in artificial neural networks that separates input representation from output production. An encoder transforms input into an internal representation, and a decoder uses that representation to generate or reconstruct an output. In sequence-to-sequence learning, the input and output may have different lengths, as in machine translation. The architecture specifies the relationship between two components rather than a particular neural-network layer type or training objective. (arxiv.org)

Structure and mathematical formulation

Let xx denote an input and z=Eθ(x)z=E_\theta(x) its encoded representation. A decoder DϕD_\phi produces an output from zz, potentially using additional information such as previously generated symbols. The parameters θ\theta and ϕ\phi are commonly learned jointly. This arrangement supports representation learning: the encoder learns features useful for the decoder’s task rather than relying solely on manually specified representations. In early recurrent sequence models, zz was a fixed-length vector; attention-based models instead expose a collection of encoder states to the decoder. (arxiv.org)

For an autoregressive output sequence y=(y1,…,yT)y=(y_1,\ldots,y_T), the conditional probability distribution is factorized as

p(y∣x)=∏t=1Tp(yt∣y<t,Eθ(x)).p(y\mid x)=\prod_{t=1}^{T}p(y_t\mid y_{<t},E_\theta(x)).

Here y<ty_{<t} contains preceding output symbols. The decoder predicts the next symbol conditioned on both the input representation and the output prefix. An end-of-sequence symbol allows generation to terminate without requiring the target length to equal the source length. This probabilistic formulation is characteristic of sequence generation, not a requirement for every encoder–decoder network. (arxiv.org)

Recurrent models and attention

Two influential implementations appeared in 2014. Cho and colleagues introduced an RNN Encoder–Decoder that encoded source phrases and decoded target phrases, using its scores within a statistical translation system. Sutskever, Vinyals, and Le demonstrated an end-to-end translation model with multilayer long short-term memory networks. These designs used recurrent neural networks to process symbols sequentially and transmit information through a fixed-dimensional representation. (arxiv.org)

A single fixed-length representation can become an information bottleneck, particularly when the input contains many details needed at different output positions. Bahdanau, Cho, and Bengio proposed an attention mechanism that allowed the decoder to obtain a different weighted combination of source representations at each generation step. The resulting soft alignment connected target predictions to relevant source positions without requiring externally supplied word alignments. (arxiv.org)

Attention therefore changes the interface between encoder and decoder: instead of receiving only one summary vector, the decoder can consult an input-dependent memory. It does not require the encoder and decoder to have identical structures. In speech transcription, for example, Listen, Attend and Spell combined an acoustic encoder with an attention-based character decoder. (arxiv.org)

Transformer implementations

The Transformer architecture, introduced in 2017, retained the encoder–decoder organization while replacing recurrence with attention and position-wise feed-forward networks. Its encoder uses self-attention to contextualize input representations. Its decoder combines masked self-attention over the output prefix with cross-attention to encoder outputs. In cross-attention, queries come from decoder states, while keys and values come from the encoder. (arxiv.org)

Multi-head attention enables several learned attention operations within each layer, while positional encoding supplies sequence-position information. Causal masking prevents a decoder position from consulting later target symbols. Because the complete target sequence is available during training, target positions can be processed through parallel computation under this mask; ordinary autoregressive generation nevertheless proceeds one output step at a time. (arxiv.org)

The terminology also distinguishes complete encoder–decoder Transformers from encoder-only and decoder-only models. T5’s architectural comparisons explicitly examine these alternatives. “Decoder” describes a component’s role and attention restrictions; it does not necessarily mean that a model reconstructs its input or implements an inverse mathematical transformation. (jmlr.org)

Training and generation

For paired input–output examples, training commonly follows supervised learning and maximum likelihood estimation. A typical loss function is the sum of negative log probabilities of correct target symbols, equivalent to token-level cross-entropy with one-hot targets. Encoder and decoder parameters receive gradients through backpropagation, including through differentiable attention operations. (arxiv.org)

During training, the decoder is often supplied with the correct preceding target symbols, a procedure called teacher forcing. During inference, those symbols are unavailable, so the model conditions on its own generated output. This discrepancy can cause prediction errors to accumulate. It is often called exposure bias and is a property of the training–generation relationship rather than of the encoder alone. (arxiv.org)

Generation also requires a search or selection procedure. Greedy decoding chooses the highest-probability next symbol; beam search retains several candidate prefixes. These procedures approximate sequence selection under the learned model and can produce different outputs from the same parameters. Architectural design, training objective, and decoding procedure are therefore separate choices. (arxiv.org)

Applications and related architectures

In natural language processing, encoder–decoder systems support translation, summarization, and question answering. T5 expresses these and other tasks in a unified text-to-text format, using task-specific text inputs and generated text outputs. The same overall architecture can consequently serve tasks that were traditionally formulated with different output interfaces. (jmlr.org)

For speech recognition, an encoder represents acoustic observations and a decoder emits linguistic symbols. In computer vision, U-Net uses a contracting path to capture context and an expanding path to produce spatially localized predictions. Connections between corresponding resolutions preserve detailed information, illustrating that encoder–decoder communication need not pass exclusively through one compressed representation. (arxiv.org)