Unicode is a standard for the consistent representation, exchange, and processing of text in computers. Maintained by the Unicode Consortium, it assigns numerical identifiers to characters from the world’s writing systems, alongside punctuation, technical symbols, and other textual elements. It also defines character properties and text-processing rules. Unicode separates character identity from its binary representation and visual appearance, allowing the same text to be exchanged between systems without relying on language-specific character sets. A version comprises a core specification, code charts, standard annexes, and machine-readable data. (unicode.org)
Origins and standardization
Before Unicode, many character encodings represented only limited repertoires. Combining text from different languages could require switching encodings, and identical numerical values could designate different characters in different systems. Unicode developed as a common framework intended to overcome these incompatibilities. Its early design emphasized a universal repertoire, logical text order, compatibility with existing encodings, and stable character identities. (unicode.org)
The Unicode Consortium was incorporated on January 3, 1991. The first volume of Unicode 1.0 appeared in October 1991, followed by the second in June 1992. Initially conceived as a 16-bit encoding, Unicode expanded its encoding architecture with version 2.0 in July 1996. Its character repertoire and encoding forms are synchronized with ISO/IEC 10646, the international Universal Coded Character Set standard. Unicode additionally specifies properties and behavior needed for consistent text processing. (unicode.org)
Characters, code points, and glyphs
A code point is a position in Unicode’s numerical code space. Code points are conventionally written as hexadecimal numbers prefixed by U+: for example, U+0041 represents LATIN CAPITAL LETTER A. The code space extends from U+0000 through U+10FFFF. Not every position represents an assigned character; some remain unassigned or serve special purposes. (unicode.org)
Unicode distinguishes an abstract character from a glyph, its visual representation. A character can have different glyphs depending on typeface and context, while several characters may combine into one displayed form. Unicode therefore does not prescribe a single font or exact drawing for every character. (unicode.org)
This distinction underlies Han unification, which assigns shared codes to many historically related Chinese characters used across East Asian writing traditions. Unification does not require their regional glyph shapes to be identical: fonts and language-sensitive rendering can preserve typographic differences. (unicode.org)
A user-perceived character is also not necessarily one code point. A base letter and accent may form a single grapheme, while Unicode’s extended grapheme-cluster rules provide practical boundaries for text selection, cursor movement, and related operations. These boundaries can encompass several code points. (unicode.org)
Encoding forms
Unicode’s principal encoding forms represent the same Unicode scalar values using different sizes of code unit. A scalar value is any Unicode code point outside the surrogate range U+D800–U+DFFF. Surrogates support UTF-16 encoding rather than representing independent characters. (unicode.org)
- UTF-8 uses one to four eight-bit code units per scalar value. It preserves the byte values of ASCII characters, making ASCII text valid UTF-8.
- UTF-16 uses one or two 16-bit code units. Values above
U+FFFFare represented by a surrogate pair. - UTF-32 uses one 32-bit code unit per scalar value. Its fixed-width representation does not mean that every user-perceived character occupies one unit. (unicode.org)
Consequently, byte count, code-unit count, code-point count, and grapheme-cluster count can differ for the same text. A byte order mark, represented by U+FEFF, can indicate byte order in UTF-16 or UTF-32. In UTF-8, where byte order is unambiguous, it functions only as an encoding signature. (unicode.org)
Normalization and equivalence
Unicode allows some text elements to have more than one representation. For example, é can be represented by the precomposed character U+00E9 or by U+0065 followed by the combining acute accent U+0301. These sequences are canonically equivalent despite having different binary representations. Unicode normalization transforms equivalent sequences into consistent forms. (unicode.org)
Four normalization forms are defined. NFD performs canonical decomposition; NFC performs canonical decomposition followed by canonical composition. NFKD and NFKC additionally apply compatibility decomposition, which can remove distinctions such as those between certain presentation variants. Compatibility normalization can therefore discard information that remains significant in some contexts. Normalization is distinct from case-insensitive matching and does not generally make visually similar characters identical. (unicode.org)
Text processing and display
The Unicode Character Database supplies properties such as general category, script, combining class, and bidirectional class. These machine-readable properties distinguish letters, numbers, punctuation, marks, and other categories, providing a common basis for implementation. (unicode.org)
Unicode includes an algorithm for bidirectional text, allowing right-to-left and left-to-right material to coexist. Characters remain stored in logical order; the bidirectional algorithm determines their display ordering. Separate specifications describe default grapheme, word, and sentence boundaries. (unicode.org)
Encoding alone does not guarantee correct display. Font coverage, shaping, and rendering remain necessary, especially for scripts with context-dependent letter forms or complex combinations. Likewise, character-code order is not inherently the appropriate alphabetical order for every language. (unicode.org)
Identifier security
Unicode’s broad repertoire includes distinct characters that can look alike. Such confusable characters can create ambiguity in identifiers, including names used by a programming language or network service. Unicode Technical Standard #39 defines mechanisms for detecting confusable strings, examining script combinations, and establishing restricted identifier profiles. Visual confusability is context-dependent and cannot be determined perfectly from character identity alone; normalization and confusable detection address different problems. (unicode.org)