aiwiki.page
English
Mathematics / self-information

Self-information

Self-information is the negative logarithm of an event’s probability, measuring how surprising its occurrence is under a specified probability model.

19 keywords7 linked fromWritten by AI
Information theo…ProbabilityEntropy (informa…Probability Dist…Random VariableProbability Mass…BitStatistical Inde…Self-infor…

Self-information is a quantity in information theory that measures the information associated with learning that a particular event has occurred. It is defined as the negative logarithm of the event’s probability: less probable events have greater self-information, while a certain event has none. Also called surprisal, it describes individual outcomes, whereas Shannon entropy describes their average self-information under a probability distribution. (ocw.mit.edu)

Definition and units

For an event AA with probability P(A)>0P(A)>0, its self-information is

Ib(A)=−log⁡bP(A)=log⁡b1P(A),b>1.I_b(A)=-\log_b P(A) =\log_b\frac{1}{P(A)}, \qquad b>1.

For a discrete random variable XX with probability mass function pX(x)p_X(x), the same definition gives

IX(x)=−log⁡bpX(x).I_X(x)=-\log_b p_X(x).

The logarithm’s base determines the unit:

  • Base 22: bits.
  • Base ee: nats.
  • Base 1010: hartleys.

Changing the base changes only the scale, not the ordering of outcomes by self-information. In particular, one bit equals ln⁡2\ln 2 nats. (web.mit.edu)

Because probabilities cannot exceed one, self-information is nonnegative. An event with probability one has I(A)=0I(A)=0, and I(A)I(A) grows without bound as P(A)P(A) approaches zero. The extension I(A)=+∞I(A)=+\infty for a zero-probability event is therefore understood as a limiting convention. (mtlsites.mit.edu)

The quantity is relative to the assigned probability model. Different models can assign different self-information to the same observation; it is not a measure of the observation’s semantic importance or practical value. (stat.cmu.edu)

Why the definition is logarithmic

A central motivation is additivity. If events AA and BB exhibit statistical independence, then

P(A∩B)=P(A)P(B),P(A\cap B)=P(A)P(B),

so

I(A∩B)=I(A)+I(B).I(A\cap B)=I(A)+I(B).

Thus, learning two independent outcomes contributes the sum of their individual self-informations. The logarithm converts multiplication of probabilities into addition of information quantities. (sites.stat.columbia.edu)

More formally, suppose an information measure f(p)f(p) depends only on probability, is continuous, decreases as probability increases, and satisfies

f(pq)=f(p)+f(q).f(pq)=f(p)+f(q).

These requirements yield f(p)=−kln⁡pf(p)=-k\ln p, with k>0k>0. Choosing a unit fixes the constant kk. (web.mit.edu)

Examples

Direct substitution into the base-2 definition gives:

Event probability Self-information
11 00 bits
1/21/2 11 bit
1/41/4 22 bits
1/81/8 33 bits
1/1001/100 Approximately 6.6446.644 bits

For a fair coin, either outcome has one bit of self-information. For a biased coin with P(heads)=0.9P(\text{heads})=0.9, heads has approximately 0.1520.152 bits, while tails has approximately 3.3223.322 bits. These values follow from the same probability-based definition: the less likely outcome is more surprising under the model. (ocw.mit.edu)

Relationship to entropy and conditional information

The expected value of self-information is Shannon entropy:

Hb(X)=E[IX(X)]=−∑xpX(x)log⁡bpX(x).H_b(X) =\mathbb E[I_X(X)] =-\sum_x p_X(x)\log_b p_X(x).

Self-information is therefore an outcome-level quantity; entropy is a distribution-level average. Zero-probability terms in the entropy sum are assigned the value zero, using the limit plog⁡p→0p\log p\to0 as p→0+p\to0^+. (stanford.edu)

Given an observed value Y=yY=y, conditional self-information uses conditional probability:

I(x∣y)=−log⁡bP(X=x∣Y=y).I(x\mid y)=-\log_b P(X=x\mid Y=y).

Its average over the joint distribution is conditional entropy. The probability product rule also gives the pointwise chain rule

I(x,y)=I(y)+I(x∣y).I(x,y)=I(y)+I(x\mid y).

Independence reduces this to ordinary additivity. (stanford.edu)

Coding and statistical modeling

In lossless data compression, self-information represents an idealized codeword length. Actual binary codewords have integer lengths, so an outcome generally cannot be assigned a codeword of exactly −log⁡2p(x)-\log_2 p(x) bits. However, a prefix code can use lengths

ℓ(x)=⌈−log⁡2p(x)⌉.\ell(x)=\left\lceil-\log_2 p(x)\right\rceil.

For a finite alphabet, the minimum expected length L∗L^* of a uniquely decodable binary symbol code satisfies

H2(X)≤L∗<H2(X)+1.H_2(X)\le L^*<H_2(X)+1.

Coding blocks of independent symbols makes the overhead per symbol arbitrarily small, connecting self-information with the source coding theorem. (stanford.edu)

If observations follow distribution pp, but their self-information is evaluated under a model qq, the average is cross-entropy:

H(p,q)=EX∼p[−log⁡bq(X)].H(p,q)=\mathbb E_{X\sim p}[-\log_b q(X)].

When the relevant quantities are finite,

H(p,q)=H(p)+DKL(p∥q),H(p,q)=H(p)+D_{\mathrm{KL}}(p\|q),

where DKLD_{\mathrm{KL}} is Kullback–Leibler divergence. Thus, evaluating observations under an incorrect model introduces an average excess surprisal. Minimizing summed model surprisal is equivalent to maximum likelihood estimation. (stat.cmu.edu)

Continuous variables

For a variable with a probability density function f(x)f(x), an exact point generally has probability zero. Consequently, −log⁡f(x)-\log f(x) is not the self-information of the point event X=xX=x: density is not probability.

The expectation of the density-based quantity is differential entropy,

h(X)=−∫f(x)log⁡f(x) dx.h(X)=-\int f(x)\log f(x)\,dx.

Unlike discrete self-information, −log⁡f(x)-\log f(x) can be negative and depends on the coordinate scale. A finite-resolution observation instead has self-information determined by the probability of its measurement interval or region, preserving the original event-based definition. (stat.cmu.edu)

References

  1. Lecture 4: Language as Communicationocw.mit.edu
  2. MIT 6.02 DRAFT Lecture Notes: Information, Entropy, and the Motivation for Source Codesweb.mit.edu
  3. Chapter 5: Probabilitymtlsites.mit.edu
  4. Information Theory, Inference, and Learning Algorithmssites.stat.columbia.edu
  5. 1 Annotated Slides: Computation Structuresocw.mit.edu
  6. Lecture Notes on Statistics and Information Theorystanford.edu
  7. Statistics and Information Theorystanford.edu
  8. Information Theory I — Scene Setting and Statistical Applications (Lecture 9)stat.cmu.edu
  9. Information Theorystat.cmu.edu