Natural language processing (NLP) is a field concerned with computational methods for analyzing, interpreting, and generating human language. It connects computer science, artificial intelligence, and linguistics. Its subjects include written text, spoken communication, and the linguistic structures underlying them. NLP encompasses both specialized tasks, such as identifying names in documents, and broader systems for translation, search, and conversation. It overlaps extensively with computational linguistics, which also investigates language through computational models. (nlp.stanford.edu)
Linguistic foundations
Language presents several interacting levels of structure. Morphology concerns the internal composition of words; syntax describes how words combine into phrases and sentences; semantics concerns meaning; and pragmatics examines interpretation in context. Processing extended discourse additionally requires tracking entities, relationships between sentences, and the organization of conversation. These levels can be modeled separately, although practical systems frequently combine information across them. (web.stanford.edu)
Ambiguity is a central difficulty. A word may have multiple senses, a sentence may permit different grammatical analyses, and a name may designate different kinds of entities. “Washington,” for example, can refer to a person, a location, or an organization in different contexts. Resolving such alternatives requires evidence beyond the isolated word, often including neighboring expressions and document context. Linguistic analysis therefore involves selecting interpretations rather than simply matching strings to fixed meanings. (stanford.edu)
Tasks and applications
NLP tasks differ in the structures they produce. Part-of-speech tagging assigns grammatical categories to words. Named entity recognition identifies spans referring to entities such as people, locations, and organizations. Both can be formulated as sequence labeling, although entity recognition must identify boundaries as well as categories. Parsing constructs representations of grammatical structure, while relation and event extraction identify connections among entities and occurrences described in text. (web.stanford.edu)
Other tasks operate at the document or interaction level. Text classification assigns categories, including topic or sentiment. Information retrieval finds material relevant to a query; question answering produces or selects answers; and summarization condenses source material. Machine translation converts content between languages. Speech-oriented systems include speech recognition, which transcribes audio, and speech synthesis, which generates spoken output. Dialogue systems combine several capabilities to interpret requests and produce responses. (web.stanford.edu)
Representations and methods
Text processing commonly begins with tokenization, the division of input into computational units. Tokens may correspond to words, word fragments, characters, or other units. Whitespace alone does not provide a universal definition of word boundaries. Subword methods create vocabularies that represent frequent strings compactly while decomposing less frequent expressions. Tokenization choices influence how linguistic information enters a model and how generated token sequences are reconstructed as text. (stanford.edu)
Rule-based approaches encode linguistic patterns and grammatical constraints explicitly. Statistical approaches instead estimate patterns from corpora. An n-gram language model, for example, approximates the probability of a token using a limited preceding context. Such models depend on observed sequence frequencies and techniques for assigning probability to unseen combinations. Statistical language modeling offers a quantitative treatment of uncertainty rather than requiring every acceptable expression to be specified in advance. (web.stanford.edu)
Neural approaches learn numerical representations and predictive functions. A word embedding represents a word as a vector, while contextual representations vary with its surrounding text. Earlier neural sequence systems commonly used recurrent or convolutional networks. The Transformer architecture, introduced in 2017, replaced recurrence in its proposed translation model with attention-based processing, enabling greater parallelization during training. Its attention mechanism combines information from different sequence positions according to learned relevance weights. (web.stanford.edu)
Pretraining and adaptation
Pretraining learns language representations from large collections of text before adaptation to particular tasks. In self-supervised learning, training targets are derived from the data itself—for example, by hiding tokens and predicting their identities. BERT uses masked-token prediction to learn bidirectional representations that incorporate left and right context. Autoregressive models instead predict successive tokens from preceding context and can generate text through repeated prediction. (aclanthology.org)
A pretrained model can undergo fine-tuning on task-specific examples. This reduces the need to design a separate architecture for every application, although suitable examples and evaluation remain important. Large language models extend language modeling to large parameter counts and training collections. They are an important approach within NLP, but the field also includes smaller statistical models, explicit linguistic analysis, and task-specific systems. (aclanthology.org)
Evaluation and limitations
Evaluation must match the task. Classification and labeling can use precision, recall, and F-score. Language models are often assessed through perplexity, which measures predictive performance on held-out text. Translation metrics such as BLEU compare generated output with reference translations using matching word sequences. These measures assess different properties: successful prediction or reference overlap does not by itself establish factual accuracy or satisfactory performance in every application. (web.stanford.edu)
Generated language can contain hallucinations—outputs that are unsupported or incorrect—although research definitions vary. Models can also reproduce social biases learned from data. Performance is uneven across languages because resources and research attention are distributed unequally. Consequently, evaluating an NLP system involves specifying its language, task, data conditions, and relevant failure modes rather than treating one benchmark score as a universal measure of competence. (aclanthology.org)