Self-information is a quantity in information theory that measures the information associated with learning that a particular event has occurred. It is defined as the negative logarithm of the event’s probability: less probable events have greater self-information, while a certain event has none. Also called surprisal, it describes individual outcomes, whereas Shannon entropy describes their average self-information under a probability distribution. (ocw.mit.edu)
Definition and units
For an event with probability , its self-information is
For a discrete random variable with probability mass function , the same definition gives
The logarithm’s base determines the unit:
- Base : bits.
- Base : nats.
- Base : hartleys.
Changing the base changes only the scale, not the ordering of outcomes by self-information. In particular, one bit equals nats. (web.mit.edu)
Because probabilities cannot exceed one, self-information is nonnegative. An event with probability one has , and grows without bound as approaches zero. The extension for a zero-probability event is therefore understood as a limiting convention. (mtlsites.mit.edu)
The quantity is relative to the assigned probability model. Different models can assign different self-information to the same observation; it is not a measure of the observation’s semantic importance or practical value. (stat.cmu.edu)
Why the definition is logarithmic
A central motivation is additivity. If events and exhibit statistical independence, then
so
Thus, learning two independent outcomes contributes the sum of their individual self-informations. The logarithm converts multiplication of probabilities into addition of information quantities. (sites.stat.columbia.edu)
More formally, suppose an information measure depends only on probability, is continuous, decreases as probability increases, and satisfies
These requirements yield , with . Choosing a unit fixes the constant . (web.mit.edu)
Examples
Direct substitution into the base-2 definition gives:
| Event probability | Self-information |
|---|---|
| bits | |
| bit | |
| bits | |
| bits | |
| Approximately bits |
For a fair coin, either outcome has one bit of self-information. For a biased coin with , heads has approximately bits, while tails has approximately bits. These values follow from the same probability-based definition: the less likely outcome is more surprising under the model. (ocw.mit.edu)
Relationship to entropy and conditional information
The expected value of self-information is Shannon entropy:
Self-information is therefore an outcome-level quantity; entropy is a distribution-level average. Zero-probability terms in the entropy sum are assigned the value zero, using the limit as . (stanford.edu)
Given an observed value , conditional self-information uses conditional probability:
Its average over the joint distribution is conditional entropy. The probability product rule also gives the pointwise chain rule
Independence reduces this to ordinary additivity. (stanford.edu)
Coding and statistical modeling
In lossless data compression, self-information represents an idealized codeword length. Actual binary codewords have integer lengths, so an outcome generally cannot be assigned a codeword of exactly bits. However, a prefix code can use lengths
For a finite alphabet, the minimum expected length of a uniquely decodable binary symbol code satisfies
Coding blocks of independent symbols makes the overhead per symbol arbitrarily small, connecting self-information with the source coding theorem. (stanford.edu)
If observations follow distribution , but their self-information is evaluated under a model , the average is cross-entropy:
When the relevant quantities are finite,
where is Kullback–Leibler divergence. Thus, evaluating observations under an incorrect model introduces an average excess surprisal. Minimizing summed model surprisal is equivalent to maximum likelihood estimation. (stat.cmu.edu)
Continuous variables
For a variable with a probability density function , an exact point generally has probability zero. Consequently, is not the self-information of the point event : density is not probability.
The expectation of the density-based quantity is differential entropy,
Unlike discrete self-information, can be negative and depends on the coordinate scale. A finite-resolution observation instead has self-information determined by the probability of its measurement interval or region, preserving the original event-based definition. (stat.cmu.edu)
References
- Lecture 4: Language as Communicationocw.mit.edu
- MIT 6.02 DRAFT Lecture Notes: Information, Entropy, and the Motivation for Source Codesweb.mit.edu
- Chapter 5: Probabilitymtlsites.mit.edu
- Information Theory, Inference, and Learning Algorithmssites.stat.columbia.edu
- 1 Annotated Slides: Computation Structuresocw.mit.edu
- Lecture Notes on Statistics and Information Theorystanford.edu
- Statistics and Information Theorystanford.edu
- Information Theory I — Scene Setting and Statistical Applications (Lecture 9)stat.cmu.edu
- Information Theorystat.cmu.edu