Perplexity is a numerical measure of uncertainty in a probability distribution or of how well a probabilistic model predicts observed data. In machine learning, particularly natural language processing, it commonly evaluates a language model by exponentiating its average negative log-likelihood on an evaluation sequence. Lower perplexity means that the model assigns higher probability, on average geometrically, to the observed tokens. It measures predictive fit rather than uncertainty in the everyday psychological sense. (web.stanford.edu)
Mathematical definition
For a discrete random variable with distribution , perplexity is the exponential of its information entropy:
When entropy is measured in bits, the equivalent expression is . The logarithm and exponential bases must match; changing both consistently leaves perplexity unchanged. (d2l.smola.org)
For observations , an autoregressive model assigns each token a conditional probability given preceding tokens. Its empirical perplexity is
Here counts evaluated prediction targets. Through the probability chain rule, this is also
Thus, perplexity is the reciprocal of the geometric mean probability assigned to the observed tokens, not the arithmetic mean of their reciprocal probabilities. It depends on both the model and the evaluated data. (web.stanford.edu)
Interpretation and examples
Perplexity can be interpreted as an effective number of equally likely alternatives. As a direct consequence of the entropy definition, a uniform distribution over outcomes has entropy and perplexity . A distribution concentrated entirely on one outcome has perplexity 1. For a distribution over possible outcomes, entropy-based perplexity lies between 1 and . (d2l.smola.org)
For an empirical sequence score, however, the upper bound need not equal vocabulary size. If a model assigns extremely small probabilities to the tokens that actually occur, perplexity can become arbitrarily large. If any evaluated token receives probability zero, its negative log-likelihood—and hence sequence perplexity—is infinite. These properties follow directly from the sequence definition. (web.stanford.edu)
For example, suppose the probabilities assigned to three observed tokens are , , and . Substitution gives
The probabilities differ at each position, but their geometric mean is . The resulting perplexity of 4 does not mean that exactly four plausible tokens existed at every position.
Cross-entropy, likelihood, and compression
Empirical perplexity is the exponential of the mean cross-entropy loss for observed token targets. Since exponentiation is strictly increasing, minimizing that mean loss also minimizes perplexity. With the same observations and normalization, minimizing it is equivalent to maximizing likelihood. (web.stanford.edu)
At the population level, for discrete distributions and ,
The Kullback–Leibler divergence term measures the additional cross-entropy associated with using instead of the true distribution . Its nonnegativity makes the source entropy a lower bound on expected cross-entropy. This population statement does not guarantee the same inequality for every finite evaluation sample. (d2l.smola.org)
The connection to information theory gives a compression interpretation: negative log probabilities describe idealized coding costs. A perplexity of 16 corresponds to four bits per evaluated symbol. Actual lossless compression also involves coding overhead and implementation details, so perplexity is not itself a measured compressed file size. (d2l.smola.org)
Evaluation and comparability
Language-model perplexity is normally reported on a held-out test set, rather than the training data, to assess prediction on unseen material. This distinction matters because overfitting can improve training fit without comparable improvement on new text. A validation set serves a different role: selecting models or training settings. (web.stanford.edu)
Meaningful comparisons require compatible text preprocessing, prediction targets, and tokenization. Word, character, and subword perplexities use different units. Two tokenizers can divide identical text into different numbers of tokens, changing both the predicted events and the normalization denominator. A lower per-token score therefore does not automatically establish superiority across tokenizers. (web.stanford.edu)
A model’s context window also affects evaluation. Splitting long text into independent chunks removes preceding context at chunk boundaries. Overlapping sliding windows retain more context; a strided window offers a computational compromise. Such evaluation choices can change reported scores even when model parameters remain unchanged. (huggingface.co)
For aggregation, the definition requires summing token negative log-likelihoods, dividing by the total number of scored tokens, and exponentiating afterward. An arithmetic average of sentence perplexities generally yields a different quantity. (huggingface.co)
Masked models and limitations
Ordinary sequence perplexity does not directly apply to masked models such as BERT, whose token predictions can use both preceding and following context. An alternative is pseudo-perplexity: mask each token in turn, sum its conditional log probability given the remaining text, normalize, and exponentiate. This uses a pseudo-likelihood rather than the left-to-right factorization of a sequence probability, so its scores are not interchangeable with autoregressive perplexities. (aclanthology.org)
Perplexity measures token prediction, not every capability of a large language model. Performance in factual question answering, reasoning, or machine translation requires additional task-specific evaluation; perplexity alone does not directly measure those outcomes. (web.stanford.edu)
References
- Speech and Language Processingweb.stanford.edu
- Perplexity of fixed-length modelshuggingface.co
- Entropy, Cross-Entropy, and KL Divergenced2l.smola.org
- Masked Language Model Scoringaclanthology.org