Chinese word segmentation is the process of identifying word boundaries in Chinese text and dividing a continuous sequence of characters into a sequence of words. It is also a fundamental task in Natural Language Processing. Written Chinese Language generally does not mark word boundaries with spaces, so segmentation cannot rely on whitespace alone. For example, “我喜欢自然语言处理” (“I like natural language processing”) can be segmented as “我/喜欢/自然/语言/处理,” although the precise granularity depends on the conventions used. Chinese word segmentation falls within the broader scope of Tokenization (natural language processing), which also includes organizing model inputs into characters, subwords, and other units that need not correspond to words in the linguistic sense. (www-nlp.stanford.edu)
Linguistic Foundations and Segmentation Ambiguity
There is no one-to-one correspondence among Chinese Characters, Morphemes, and words. Characters are units of writing, morphemes are units of word formation, and words may consist of one or more characters. Chinese word segmentation therefore involves both character combinations and Morphology (Linguistics) and lexical structure. Different annotation conventions may treat compounds, derived forms, and fixed expressions differently. (aclanthology.org)
The main challenges include resolving ambiguity and identifying out-of-vocabulary words. Ambiguity can arise when different candidate words overlap at the same character position, or when a string of characters can either form a single word or be divided into several words. For example, “和尚” can mean “monk,” but in certain sentences its characters may belong to two separate units, “和/尚……” (“and / still…”). Dictionary lookup alone cannot determine the segmentation; context is also needed. (www-nlp.stanford.edu)
An out-of-vocabulary word is a word that is absent from a particular dictionary or training corpus, not necessarily a newly coined word. Personal names, place names, organization names, and technical terms all pose challenges of this kind. Models can use character features and context to identify these units. Supplementary dictionaries can also provide useful information, but they cannot eliminate all ambiguity. (aclanthology.org)
Main Methods
Dictionary- and rule-based methods locate boundaries using word lists and segmentation rules. Forward or backward maximum matching starts at one end of the text and preferentially selects the longest matching word in the dictionary. These methods are straightforward, but choosing the longest local match does not necessarily produce the most appropriate segmentation for the whole sentence, and expressions outside the word list require additional handling. Another approach organizes candidate segmentations into a word graph and uses Dynamic programming to compare paths, rather than making irrevocable choices at each step. (aclanthology.org)
Statistical methods learn segmentation patterns from annotated corpora. Supervised learning commonly recasts segmentation as sequence labeling: each character receives a label indicating its position within a word, and the labels are then converted into a word sequence. The common BMES labels denote the beginning, middle, and end of a word and a single-character word, respectively. Binary labels indicating whether a word boundary is present can also be used. Hidden Markov models, maximum entropy models, and conditional random fields have all been used for this task. Conditional random fields can incorporate features drawn from neighboring characters, character combinations, and dictionaries. (aclanthology.org)
Neural network methods use Deep Learning to learn character representations and contextual features, reducing reliance on manual Feature engineering. Researchers have used Long short-term memory networks as well as pretrained encoders such as BERT (language model), followed by Fine-tuning (deep learning) to predict word boundaries. Models that support multiple annotation conventions can incorporate convention identifiers and Multi-task Learning, allowing a single model to accommodate different segmentation requirements. (arxiv.org)
In addition, Unsupervised learning methods attempt to infer word units from recurring patterns and statistical structure in text when manually annotated word boundaries are unavailable. For example, a segmental Language model can use candidate word segments as its generative units, although the automatically obtained segments do not necessarily align fully with human annotation conventions. (arxiv.org)
Annotation Conventions and Evaluation
Chinese word segmentation has no single granularity suitable for every purpose. Datasets differ in their conventions for compound structures, proper names, numerical expressions, and other phenomena. Different segmentations of the same text may therefore reflect different task conventions rather than simply algorithmic errors. Whether the Training data and evaluation data follow the same conventions affects how results should be interpreted. (nlp.stanford.edu)
The Second International Chinese Word Segmentation Bakeoff, held in 2005, used four corpora supplied by Academia Sinica, City University of Hong Kong, Peking University, and Microsoft Research. It included open and closed tracks that distinguished between different levels of access to additional resources. These datasets subsequently became important benchmarks for Chinese word segmentation research. (aclanthology.org)
Evaluation usually takes manually annotated words as the reference: a predicted word is counted as correct only if both its start and end positions match the reference. If the number of correct words is (C), the number of predicted words is (N_p), and the number of reference words is (N_g), then precision is (P=C/N_p), recall is (R=C/N_g), and the F1 score is (2PR/(P+R)). Evaluations also commonly report out-of-vocabulary word recall separately. Reporting only an overall score can conceal a system’s weaknesses in identifying low-frequency words. (aclanthology.org)
Applications and the Relationship to Tokenization
Chinese word segmentation can support tasks such as information retrieval, Machine translation, and part-of-speech tagging. Word boundaries determine the units used in subsequent processing, while segmentation consistency and granularity also affect application performance. A high score on a segmentation benchmark therefore does not automatically imply the best performance on every downstream task. (nlp.stanford.edu)
Character-level models mean that explicit word segmentation is no longer a necessary preprocessing step for every Chinese-language system. For example, the original Chinese BERT primarily builds representations at the character level, but research has found that it can still encode word-structure information internally. This calls for a distinction between two goals: Chinese word segmentation identifies word boundaries according to established conventions, whereas model tokenization constructs the input units needed for computation. A single word may correspond to multiple tokens, and the two processes need not produce the same segmentation. (aclanthology.org)