A grapheme is a minimal functional or distinctive unit of a writing system. It is an abstract unit rather than a particular handwritten mark or printed shape: different visual forms may realize the same grapheme. In linguistics, the concept helps describe how writing distinguishes words and represents language. Its precise definition varies, particularly over whether units should be identified through written contrasts or their relationship to speech. In computing, the related term grapheme cluster denotes an algorithmically identified text segment approximating a user-perceived character, not necessarily a linguistic grapheme. (unicode.org)
Distinctiveness and abstraction
A common way to identify graphemes is to examine contrasts between written words. In English, big and dig demonstrate that ‹b› and ‹d› are distinct units: substituting one for the other produces a different word. This resembles the use of minimal pairs in phonology to establish distinct phonemes. The written contrast, however, concerns the organization of writing rather than simply the sounds associated with its symbols. (unicode.org)
The distinction between an abstract unit and its realization is essential. A Roman-style lowercase a and an italic lowercase a do not normally distinguish English words. They can therefore represent the same grapheme despite their different shapes. Such variants are called allographs. A glyph, in digital typography, is a visual form selected to display one or more characters; it belongs to the representation of text rather than automatically defining a distinctive linguistic unit. (unicode.org)
Graphemes are frequently enclosed in angle brackets, as in ‹a›, while phonemes appear between slashes, as in /a/. This notation keeps written units separate from their phonological interpretation. The distinction is especially useful when a written sequence corresponds to one sound or when a written form has several pronunciations. (meletis.at)
Competing linguistic definitions
Two influential approaches are often termed referential and analogical. The referential approach defines graphemes through their correspondence with phonemes. Under this interpretation, a sequence of letters may constitute one grapheme: German ‹sch›, for example, represents /ʃ/. Different spellings associated with the same phoneme may also be grouped as variants of one grapheme. This approach makes the relation between writing and speech central to the analysis. (meletis.at)
The analogical approach instead treats writing as a system whose distinctive units can be established by methods analogous to those used for speech. Written minimal pairs and distributional patterns provide evidence. A sequence representing one phoneme need not therefore be one grapheme. The two approaches can produce different inventories for the same writing system, and neither terminology alone resolves questions about minimality or the status of individual letters. (meletis.at)
Broader proposals seek a definition applicable beyond the alphabet. One proposal combines three criteria: a grapheme distinguishes lexical meaning, has linguistic value, and cannot be decomposed into smaller units that themselves qualify as graphemes. Its linguistic value may involve a phoneme, syllable, or morpheme. This is a proposed comparative framework, not a universally accepted definition. (meletis.at)
Variation across writing systems
The identification of graphemes depends on a system’s structure, not simply on visible complexity. Chinese characters may contain several components while functioning as single graphemes. In the proposed framework above, ‹河›, representing the morpheme meaning “river,” qualifies as a grapheme, whereas its reduced component ‹氵› does not independently represent a linguistic unit in that occurrence. Japanese kana provide examples of units corresponding to a mora, a phonological timing unit. (meletis.at)
Spatial arrangement can complicate segmentation. Korean Hangul places component letters together in syllable blocks; a linguistic analysis can identify several graphemes within one block. Conversely, some accounts of Devanagari treat a consonant-and-vowel configuration as one grapheme. Consequently, a visually bounded group is not a language-independent test of grapheme status. (meletis.at)
The treatment of punctuation, digits, and other symbols also varies. Punctuation can distinguish sentence functions and convey information about syntax or pragmatics, but does not necessarily correspond to a phoneme or morpheme. Whether such signs count as graphemes depends partly on how narrowly linguistic value is defined. (meletis.at)
Graphemes and digital encoding
Unicode distinguishes encoded characters from their displayed forms. A code point identifies a position in the encoding space, not necessarily one complete user-perceived character. For example, å can be encoded as a precomposed character or as a base letter followed by a combining ring. These representations can express the same written unit while containing different numbers of code points. (unicode.org)
Unicode normalization provides standardized transformations between equivalent encoded representations. Canonical decomposition separates eligible precomposed characters into components; canonical composition recombines eligible sequences. Normalization concerns representation and equivalence, rather than deciding which units are linguistically graphemes. (unicode.org)
Unicode’s grapheme clusters supply practical segmentation units. Default extended grapheme clusters group sequences such as base letters with combining marks and Hangul syllable sequences. They approximate user-perceived characters without requiring language-specific analysis. A linguistic multi-letter unit may nevertheless remain several default clusters, while one cluster may contain multiple linguistic graphemes. (unicode.org)
These boundaries support cursor movement, text selection, deletion, and character counting. They are distinct from glyph boundaries: a font may display fi as one ligature, although the underlying sequence remains two default grapheme clusters. Segmentation may be tailored for particular languages or operations, and a cluster count measures perceived text units rather than storage size. (unicode.org)