A large language model (LLM) is a language model built with an artificial neural network and trained on extensive collections of text to predict, represent, or generate linguistic sequences. LLMs are used in natural language processing for tasks including question answering, summarization, translation, and dialogue. “Large” refers to substantial model and training scale rather than a universally agreed parameter threshold. The term commonly denotes broadly capable pretrained models, especially generative systems, rather than a particular application or chatbot interface. (arxiv.org)
Historical development
LLMs developed from statistical language modeling and earlier neural approaches. Their predecessors included models that estimated word-sequence probabilities from limited contexts and networks that learned continuous linguistic representations. Increasingly, pretraining on broad text collections became a way to obtain reusable representations instead of training a separate model from scratch for every task. (arxiv.org)
The Transformer architecture, introduced in the 2017 paper Attention Is All You Need, replaced recurrent sequence processing with attention-based computation. Its ability to parallelize training helped make large-scale language-model development practical. In 2018, BERT demonstrated bidirectional pretraining for language understanding, while the generative pretrained Transformer family developed autoregressive text generation. The 2020 GPT-3 study showed that a 175-billion-parameter model could perform numerous tasks from instructions and examples without task-specific parameter updates. (arxiv.org)
Architecture and text representation
Text is converted into discrete units through tokenization. Tokens may represent words, word fragments, punctuation, or byte-level units. These identifiers are mapped to numerical vectors and processed through successive network layers. Transformer layers use self-attention to combine information from different positions, while positional representations supply information about sequence order. A model’s context window limits how much input and generated text it can process together in a given computation. (arxiv.org)
Architectures differ in how they access context. Encoder models such as BERT construct representations using surrounding text and are commonly adapted for classification or extraction. Decoder-only models generate sequences by predicting successive tokens from preceding tokens. Encoder–decoder models separately process an input and produce an output conditioned on it. These distinctions influence training objectives and suitable applications; not every pretrained language model is primarily designed for open-ended conversation. (arxiv.org)
Pretraining and adaptation
Pretraining typically uses self-supervised learning, in which prediction targets are derived from the data itself. An autoregressive model learns to predict the next token; a masked language model learns to reconstruct selected missing tokens. A common loss function is cross-entropy, which penalizes predictions assigning low probability to the observed target. Training adjusts parameters through backpropagation and numerical optimization. (arxiv.org)
Training data may include web pages, books, articles, and source code. Selection, filtering, deduplication, and the proportions of different data sources affect model behavior. Neural scaling laws describe empirical relationships among predictive performance, model size, data volume, and computation. The 2022 Chinchilla study demonstrated that allocating a fixed training budget between parameters and training tokens matters: increasing parameter count alone is not necessarily the best use of computation. (arxiv.org)
After pretraining, fine-tuning can adapt a model to particular domains or interaction styles. Instruction tuning uses examples of requests and desired responses. Reinforcement learning from human feedback can further shape behavior using preferences over outputs. These stages distinguish a general text-prediction model from an instruction-following assistant, but they do not eliminate factual or behavioral errors. (arxiv.org)
Generation and applications
During inference, an autoregressive model produces a probability distribution over possible next tokens. A decoding procedure selects a token, appends it to the sequence, and repeats. Greedy selection and probabilistic sampling can produce different outputs from the same prompt; sampling settings affect variability, not a direct guarantee of truthfulness. Ordinary inference uses already learned parameters rather than retraining the model for each request. (arxiv.org)
In-context learning allows instructions or demonstrations in a prompt to influence task performance without parameter updates. Applications include machine translation, document summarization, text classification, question answering, and code generation. Performance varies across tasks, languages, and prompt formulations, so capability in one setting does not establish equal capability elsewhere. (arxiv.org)
An LLM can also operate within a larger system. Retrieval-augmented generation supplies relevant passages from an external collection, allowing generation to use information beyond what is encoded in model parameters. Retrieval can improve factual specificity and support access to updateable knowledge, although the system still depends on the quality of retrieved evidence and its use by the generator. (arxiv.org)
Evaluation and limitations
Evaluation combines language-prediction measures with task benchmarks and human judgments. Perplexity measures predictive fit to text, but does not by itself establish usefulness, factual accuracy, or safety. Frameworks such as HELM evaluate multiple dimensions, including accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency. Comparisons also depend on prompts, decoding procedures, and evaluation conditions. (arxiv.org)
Hallucination describes generated content that is factually incorrect or unsupported by the relevant evidence. Fluent language can therefore coexist with invented details or erroneous reasoning. Models can also reproduce undesirable patterns in training material. Benchmark interpretation requires attention to data leakage, including overlap between training material and evaluation examples. Human-feedback training can reduce some unwanted outputs, but neither model scale nor polished presentation establishes reliability. (arxiv.org)