A bit, short for binary digit, is a unit of digital data that takes one of two values, conventionally written 0 and 1. In computer science, bits provide the basic representation used by digital computers. In information theory, the bit is also a unit for measuring information and uncertainty using base-two logarithms. These meanings are related but distinct: a stored binary digit does not necessarily contain one bit of unpredictable information. (cs.cornell.edu)
Origin and terminology
Claude Shannon used the term in his 1948 paper A Mathematical Theory of Communication, published in two parts in July and October in the Bell System Technical Journal. He credited J. W. Tukey with suggesting the abbreviated word. Shannon explained that a device with two stable positions, such as a relay or flip-flop circuit, can store one bit; a collection of such devices has possible configurations. (cs.yale.edu)
The same term consequently describes both an individual binary symbol and a logarithmic unit of information. A bit value might signify a number, a logical condition, or part of a larger encoded message. Its interpretation depends on the conventions of the system using it, rather than on the symbols 0 and 1 alone. (cs.cornell.edu)
Binary representation
In the binary number system, each position has a weight equal to a power of two. An unsigned sequence of bits represents an integer from 0 through . For example, the sequence represents in decimal notation. Four bits permit 16 different patterns, while eight permit 256. These are consequences of assigning two possible values to every position. (cs.cornell.edu)
The numerical meaning of a pattern depends on its encoding. Under two’s complement, the most significant bit of an -bit signed integer has weight , rather than . Thus, an eight-bit pattern can represent a different value when interpreted as signed rather than unsigned. The bits themselves do not change; the interpretation does. (cs.cornell.edu)
Physical implementation and logical operations
A bit is an abstraction, not a particular material object. Electronic circuits map physical conditions to logical values. Transistors act as controlled switches, and combinations of switches implement operations on binary inputs. This separation between physical behavior and logical interpretation allows circuits to be described without specifying every microscopic detail of their construction. (courses.cs.cornell.edu)
The operations are described using Boolean algebra and implemented by logic gates. AND, OR, and NOT combine or invert binary values, while exclusive OR produces 1 when its two inputs differ. Networks of gates can implement arithmetic, including addition with carry bits. Larger computational structures build on these elementary operations. (cs.cornell.edu)
A physical storage cell need not correspond to exactly one bit. Flash memory can distinguish multiple stored states: two bits per cell require four distinguishable states, three bits require eight, and four bits require sixteen. Consequently, counting physical cells and counting logical bits are different ways of describing storage. (micron.com)
Bits, bytes, and prefixes
A byte conventionally contains eight bits; a four-bit group is called a nibble. Bits count binary digits, whereas bytes group them into larger units. Confusing the two introduces a factor-of-eight error when converting a capacity or transfer rate. The symbol B denotes a byte; bit explicitly denotes a bit. (cs.cornell.edu)
Decimal prefixes retain their SI meanings: a kilobit is bits and a megabit is bits. Binary prefixes explicitly denote powers of two: a kibibit is bits and a mebibit is bits. Similarly, one megabyte is 1,000,000 bytes, whereas one mebibyte is 1,048,576 bytes. (physics.nist.gov)
Information content
The information associated with an event of probability is its self-information,
An outcome with probability has one bit of self-information; an unlikely outcome can have more. This quantity measures surprise relative to a probability model, not the number of binary digits physically allocated to record the outcome. (ocw.mit.edu)
For a binary random variable with probabilities and , its information entropy is
with zero-probability terms taken as zero. Entropy equals one bit when both outcomes are equally likely and falls to zero when the outcome is certain. A predictable binary digit therefore occupies a bit position while contributing no uncertainty. (ocw.mit.edu)
Lossless data compression exploits unequal probabilities and predictable structure to shorten representations. For an independent, identically distributed source, the source coding theorem relates achievable average coding length to source entropy. Information measured in bits is therefore not simply synonymous with file size. (ocw.mit.edu)
Errors and quantum information
Stored or transmitted bits may be corrupted. An error-correcting code introduces structured redundancy so that specified error patterns can be detected or corrected. A single parity bit detects any odd number of bit flips but cannot detect every even-numbered error pattern or, by itself, locate a flipped bit. (ocw.mit.edu)
A qubit differs from a classical bit because its state can be a quantum superposition of the basis states and . Measurement in that basis yields a classical result, 0 or 1, with probabilities determined by the state’s amplitudes. Superposition is not a third classical bit value; it belongs to a different mathematical description of information. (learning.quantum.ibm.com)