Tokenization in natural language processing is the conversion of text into an ordered sequence of units called tokens. Depending on the system, tokens may represent words, punctuation, subword fragments, characters, or bytes. In neural systems, these units are usually mapped to numerical identifiers in a vocabulary. Tokenization establishes the representation on which subsequent computation operates; it does not necessarily recover linguistic words or meaningful components. Different tokenizers can therefore produce different sequences from identical text. (huggingface.co)
Tokens and linguistic boundaries
A token is a computational unit rather than a universal unit of language. For illustration, a word-oriented tokenizer might divide “Cats sleep.” into Cats, sleep, and .. Another tokenizer could divide Cats into smaller pieces. Such choices depend on the vocabulary, segmentation rules, and intended application. Subword pieces need not coincide with morphemes, the units studied in morphology: statistically useful fragments can cross linguistic boundaries or have no independent meaning. (huggingface.co)
Whitespace offers a useful starting point in English, but punctuation and contractions complicate word segmentation. Other writing systems do not consistently mark word boundaries with spaces. Chinese word segmentation, for example, requires decisions about how character sequences form words. Unicode’s default boundary rules provide a general framework, but explicitly allow language-specific tailoring rather than claiming to resolve every ambiguity. (tensorflow.org)
“Character” also requires qualification. A Unicode code point is not necessarily a complete user-perceived grapheme: combining marks and some emoji sequences contain several code points. Character-level tokenization must specify which representation it uses. Byte-level tokenization operates below this level, on encoded byte sequences rather than displayed characters. (unicode.org)
Major approaches
Word-level tokenization uses words or word-like units, often separating punctuation. In fixed-vocabulary models, this can require large vocabularies and leave unfamiliar spellings, names, or inflected forms without dedicated entries. Such inputs may be represented by an unknown-token marker. Character-level tokenization uses smaller units and a smaller vocabulary, but generally produces longer sequences. Its coverage still depends on which characters the vocabulary contains. (huggingface.co)
Subword tokenization occupies an intermediate position. Frequent words may remain intact, while rarer forms are represented as sequences of reusable fragments. This approach helps represent unfamiliar words without assigning every possible word its own entry. A influential 2016 study by Rico Sennrich, Barry Haddow, and Alexandra Birch demonstrated subword-based handling of rare and unknown words in neural machine translation. (research.ed.ac.uk)
Three prominent methods differ in how they construct or apply their vocabularies:
- Byte-pair encoding (BPE) begins with small units and repeatedly merges frequent adjacent pairs. Tokenization subsequently applies the learned merge rules. Byte-level variants start from byte values, allowing text coverage without requiring a separate vocabulary entry for every Unicode character. (huggingface.co)
- WordPiece, used in BERT, segments words using a learned subword vocabulary. Its standard segmentation procedure proceeds left to right, selecting the longest available matching piece at each position. Implementations commonly distinguish word-initial pieces from continuation pieces. (arxiv.org)
- Unigram subword models assign a probability to each vocabulary piece and evaluate alternative segmentations. Training progressively removes less useful candidates; decoding typically selects the highest-probability segmentation. The model can also sample alternatives instead of always choosing one fixed sequence. (huggingface.co)
SentencePiece, introduced in 2018, is a tokenizer and detokenizer framework rather than a synonym for one segmentation algorithm. It supports learning subword representations directly from raw sentences, avoiding a requirement for language-specific word segmentation beforehand. (aclanthology.org)
The tokenization pipeline
A practical tokenizer often contains several stages. Normalization may standardize Unicode representations, change letter case, or remove accents. Pre-tokenization divides the input into preliminary spans, such as whitespace-delimited strings. A learned segmentation algorithm then produces vocabulary tokens and their identifiers. Post-processing can insert special tokens that identify sequence boundaries or separate paired inputs. Some implementations also preserve offsets linking tokens to the original text. (huggingface.co)
Learning a tokenizer is distinct from learning a language model. The tokenizer acquires vocabulary entries, merge rules, or segmentation probabilities from training data. The downstream model learns how to process or predict the resulting sequences. Embedding layers associate token identifiers with learned vectors; the identifiers themselves are labels, not numerical measures of meaning. (huggingface.co)
Detokenization reconstructs text from tokens, but exact reversibility depends on the complete pipeline. Lowercasing or other lossy normalization can prevent recovery of the original spelling even when the underlying subword segmentation is reversible. (tensorflow.org)
Consequences for language models
Vocabulary size and sequence length create competing computational demands. Larger vocabularies require larger embedding tables, whereas finer segmentation generates more positions to process. Tokenization also affects how much text fits inside a fixed context window. Studies of multilingual large language models have found that comparable content can require substantially different token counts across languages, influencing effective context capacity and token-based costs. (tensorflow.org)
Segmentation need not remain deterministic during model training. Subword regularization samples multiple segmentations of the same text, functioning as data augmentation and regularization. Kudo’s 2018 experiments reported improvements in translation, particularly in low-resource and out-of-domain settings. These findings concern the tested configurations, not a guarantee that sampled segmentation benefits every model or task. (aclanthology.org)