The bag-of-words model is a representation of text used in natural language processing, information retrieval, and machine learning. It describes a document through the words it contains and their occurrence counts, rather than their sequence. The name “bag” indicates that repetitions are retained but ordering is discarded. This converts variable-length documents into numerical features that can support comparison, classification, and other computational tasks. It is a representation scheme, not a complete predictive algorithm. (nlp.stanford.edu)
Mathematical representation
Let a vocabulary be an ordered list of distinct terms, (V=(t_1,\ldots,t_m)). A document (d) is represented by
[ \mathbf{x}(d)=(c(t_1,d),\ldots,c(t_m,d)), ]
where (c(t_i,d)) counts occurrences of term (t_i). Each vocabulary item defines a coordinate in a vector space. Documents therefore share the same dimensions even when their lengths and contents differ. (nlp.stanford.edu)
For example, choose the vocabulary (cats, chase, dogs). Then:
| Document | Count vector |
|---|---|
| “cats chase dogs” | ((1,1,1)) |
| “dogs chase cats” | ((1,1,1)) |
| “cats chase cats” | ((2,1,0)) |
The first two sentences have different meanings but identical representations. The third shows that counts distinguish repeated words from single occurrences.
Stacking vectors as rows produces a document–term matrix, with documents as rows and vocabulary terms as columns. Because an individual document normally contains only a small fraction of the full vocabulary, implementations commonly use a sparse matrix, storing nonzero entries rather than every coordinate. (scikit-learn.org)
Vocabulary construction and preprocessing
Construction begins with tokenization: dividing text into units that become candidate vocabulary entries. Although called “words,” these units depend on the tokenizer and preprocessing rules. Punctuation, numbers, capitalization, and word boundaries can affect which features are created. A consistent token-to-coordinate mapping is necessary when transforming subsequent documents. (scikit-learn.org)
Preprocessing may merge related surface forms. Stemming applies rules that remove or modify word endings, whereas lemmatization uses vocabulary and morphological analysis to identify base forms. Both can reduce vocabulary fragmentation, but they are separate processing choices rather than defining properties of bag-of-words representation. (nlp.stanford.edu)
A system may also remove stop words, such as frequent function words judged unhelpful for a particular retrieval task. Such removal changes the representation: terms excluded from the vocabulary contribute no information to later comparisons or predictions. Stop-word selection is therefore task-dependent, not an obligatory part of the model. (nlp.stanford.edu)
Counts and term weighting
Raw counts are the simplest feature values, but several alternatives preserve the underlying unordered representation. A binary variant records only presence or absence. Relative-frequency and logarithmic transformations modify how repeated occurrences affect a document’s features. These choices distinguish the vocabulary-based representation from the weighting scheme applied to it. (scikit-learn.org)
Term frequency–inverse document frequency, usually abbreviated TF–IDF, combines a document-level occurrence measure with a collection-level rarity measure. A common formulation is
[ w(t,d)=\operatorname{tf}(t,d)\log\frac{N}{\operatorname{df}(t)}, ]
where (N) is the number of documents and (\operatorname{df}(t)) is the number containing term (t). Terms appearing in relatively few documents receive larger inverse-document-frequency weights. Document frequency differs from the total number of occurrences across the collection. Implementations may vary the scaling and smoothing conventions. (nlp.stanford.edu)
Weighted vectors can be compared using cosine similarity:
[ \operatorname{sim}(\mathbf{x},\mathbf{y}) =\frac{\mathbf{x}\cdot\mathbf{y}} {|\mathbf{x}|_2|\mathbf{y}|_2}. ]
For nonzero vectors, this compares their directions rather than their unnormalized magnitudes. The numerator is their inner product. Cosine normalization helps prevent longer documents from receiving larger similarity scores merely because they contain more words. (nlp.stanford.edu)
Applications and statistical assumptions
In supervised learning, bag-of-words features support document classification. In unsupervised learning, they can support clustering and topic discovery. Retrieval systems likewise compare vocabulary-based representations of queries and documents. These applications use the representation as input to a separate scoring or learning procedure. (scikit-learn.org)
The naive Bayes classifier illustrates an important distinction between representation and statistical assumptions. Multinomial text classification uses occurrence information; its formulation assumes that term probabilities do not depend on position and that occurrences are conditionally independent given the class. Bernoulli text classification instead models term presence and absence. However, adopting bag-of-words features does not itself require every downstream classifier to assume conditional independence between features. (nlp.stanford.edu)
Limitations and extensions
Discarding order removes information carried by syntax and sentence structure. The example above shows that a model can fail to distinguish who performs an action from who receives it. Individual-word counts also fail to represent multiword expressions directly, and spelling variations or derived forms remain separate features unless preprocessing connects them. (scikit-learn.org)
A bag-of-n-grams representation extends the approach by counting consecutive sequences. Bigrams can preserve local expressions such as “not useful,” while character n-grams capture overlapping portions of words and can provide robustness to spelling variation. Nevertheless, the resulting document remains an unordered collection of features: local order is encoded within each n-gram, but most global structure is lost. (scikit-learn.org)
Evaluation and implementation
When evaluating a learned system, vocabulary construction and collection-derived weights belong to preprocessing fitted on training data. Applying the fitted transformation to held-out documents preserves the feature space without learning from the test set. Using held-out information during preprocessing can introduce data leakage. During cross-validation, preprocessing must therefore be fitted separately within each training fold; pipelines can enforce this separation. (scikit-learn.org)
References
- Term frequency and weightingnlp.stanford.edu
- 2. Feature extractionscikit-learn.org
- Stemming and lemmatizationnlp.stanford.edu
- Inverse document frequencynlp.stanford.edu
- Tf-idf weightingnlp.stanford.edu
- Term weighting summarynlp.stanford.edu
- Dot productsnlp.stanford.edu
- Properties of Naive Bayesnlp.stanford.edu
- Common pitfalls and recommended practicesscikit-learn.org