A generative pre-trained transformer (GPT) is a type of language model that uses the transformer architecture to learn patterns in text and generate continuations. Its defining approach combines generative pre-training on large text collections with subsequent adaptation or prompting for particular tasks. The name is associated especially with OpenAI’s GPT model family, beginning with the model introduced in 2018. GPT models have become an important approach to large language models and generative artificial intelligence. (openai.com)
Meaning and historical development
The three components of the name describe complementary features. Generative refers to modeling sequences so that new continuations can be produced. Pre-trained indicates that an initial, broadly applicable training stage precedes task-specific adaptation. Transformer identifies the neural-network architecture. The original GPT work applied transfer learning to natural language processing: a language model learned from unlabeled text, then underwent fine-tuning using labeled examples for individual tasks. (cdn.openai.com)
The architectural foundation appeared in the 2017 paper Attention Is All You Need. Its Transformer used attention rather than recurrence as its central sequence-processing mechanism. OpenAI announced its initial generative pre-training research on June 11, 2018, demonstrating improvements on tasks including textual entailment, question answering, and document classification. (arxiv.org)
GPT-2, announced on February 14, 2019, extended this approach to a model with approximately 1.5 billion parameters trained on a collection of eight million web pages. Its research emphasized performing tasks through text conditioning without task-specific training. The 2020 GPT-3 paper described a 175-billion-parameter model and systematically evaluated task performance when instructions and examples were supplied in the input. The 2023 GPT-4 technical report extended the family to a multimodal model accepting image and text inputs and producing text outputs; it did not disclose model size or detailed architecture. These are milestones rather than an exhaustive model inventory. (openai.com)
Architecture and sequence representation
The original GPT and GPT-2 use a decoder-only Transformer. Unlike the original Transformer’s encoder–decoder architecture, this arrangement processes a single sequence without a separate encoder. Text first undergoes tokenization, producing discrete units that may represent words, word fragments, or other character sequences. GPT-2 uses a byte-level form of byte-pair encoding, allowing text to be represented without relying on a fixed vocabulary of complete words. (cdn.openai.com)
Tokens are mapped to learned vectors and combined with information about their sequence positions. Stacked blocks apply multi-head attention, feed-forward transformations, residual connections, and normalization. Within each block, causal self-attention prevents a position from attending to later positions. Consequently, the representation used to predict a token depends on preceding context rather than on the token’s future continuation. Positional information supplies ordering that attention alone does not inherently provide. (arxiv.org)
This differs from the original BERT, which learns bidirectional representations by predicting masked tokens using context on both sides. GPT’s causal objective directly supports left-to-right generation, whereas BERT’s original objective emphasizes contextual representation rather than sequential continuation. (arxiv.org)
Pre-training objective
GPT pre-training typically uses self-supervised learning: prediction targets come from the text itself rather than from separately supplied human labels. For a token sequence , the autoregressive model factorizes its probability as
Training maximizes the likelihood of observed sequences, or equivalently minimizes their negative log-likelihood. This is commonly expressed as a cross-entropy loss function over next-token predictions. (cdn.openai.com)
The training data determines which linguistic patterns, topics, and styles are available to learn. Model parameters are adjusted through gradient-based optimization using backpropagation. Causal masking permits predictions for many positions to be computed together during training, although ordinary generation remains sequential because each newly generated token becomes context for the next. (cdn.openai.com)
Research on neural scaling laws documents empirical relationships between language-model loss, parameter count, dataset size, and training computation. These relationships concern measured predictive performance under particular experimental conditions; they do not establish that increasing model size guarantees factual correctness or competence on every task. (arxiv.org)
Adaptation and generation
A pre-trained GPT can be adapted by supervised learning on task examples. Alternatively, in-context learning supplies instructions or demonstrations within the input while leaving model parameters unchanged. GPT-3’s evaluations distinguished zero-shot, one-shot, and few-shot settings according to the number of demonstrations provided. Prompt-based task specification therefore differs from fine-tuning, which updates model weights. (arxiv.org)
Instruction-following models may undergo additional training on human-written demonstrations and preferences. InstructGPT combined supervised fine-tuning with a reward model trained on ranked outputs and reinforcement learning from human feedback. These post-training procedures modify behavior but are not defining requirements of generative pre-training itself. (arxiv.org)
During generation, a prompt conditions the distribution over the next token. A decoding procedure selects or samples a token, appends it to the sequence, and repeats. Different sampling choices can produce different continuations from the same prompt. Text-conditioned applications include question answering, summarization, machine translation, and dialogue. (cdn.openai.com)
Limitations and evaluation
GPT output can be fluent while containing unsupported or false statements, a failure commonly called AI hallucination. The GPT-4 report documents hallucinations, reasoning errors, and sensitivity to inputs despite improved benchmark performance. Human-feedback training can reduce some undesirable outputs without eliminating them. (arxiv.org)
Evaluation must also distinguish genuine task performance from exposure to test material during pre-training. The GPT-3 study investigated overlap between its training corpus and evaluation datasets and identified contamination as a methodological concern. Reported results therefore depend on dataset construction, prompting conditions, and the measures used, rather than constituting a single universal measure of model capability. (arxiv.org)