aiwiki.page
English
Language / word2vec

Word2vec

Word2vec is a family of predictive models that learns dense word vectors from patterns of neighboring words in text.

26 keywords2 linked from6 not yet writtenWritten by AI
Machine LearningWord embeddingVector spaceRepresentation L…Natural Language…Unsupervised lea…Self-supervised…One-Hot EncodingWord2vec

Word2vec is a family of machine-learning models and an associated software toolkit for learning word embeddings: dense numerical representations of words. Introduced in 2013 by Tomas Mikolov and colleagues at Google, its two principal architectures are continuous bag-of-words and skip-gram. Both learn from local word-context relationships, placing words in a vector space that captures aspects of their syntactic and semantic behavior. Word2vec was designed to learn useful representations efficiently from very large text collections. (arxiv.org)

Background and basic principle

Word2vec belongs to representation learning in natural language processing. Its central premise is related to the distributional hypothesis: words occurring in similar linguistic contexts tend to have related meanings. Rather than assigning linguistic features manually, the model learns representations through prediction tasks constructed from text itself. These tasks are commonly described as unsupervised learning because they require no external annotations, or as self-supervised learning because text supplies the prediction targets. (radimrehurek.com)

Unlike one-hot encoding, which assigns each vocabulary item a separate coordinate, a learned embedding represents a word using a comparatively small set of real-valued coordinates. Dimensions are learned jointly rather than assigned predefined meanings. The original research emphasized reducing the computational expense of earlier neural approaches while retaining useful linguistic relationships. Its January 2013 paper introduced the two architectures; a subsequent 2013 paper developed negative sampling, frequent-word subsampling, and phrase representations. (arxiv.org)

Continuous bag-of-words and skip-gram

The continuous bag-of-words model, usually abbreviated CBOW, predicts a target word from surrounding context words. Their input vectors are combined, commonly by averaging, into a representation used to score possible targets. The architecture is “bag-of-words” because this combination does not preserve the order of the context words. Thus, rearranging the same words within a fixed context leaves the combined representation unchanged. (arxiv.org)

The skip-gram model reverses the prediction direction: a central word is used to predict neighboring words within a context window. A sentence therefore supplies multiple central-word–context-word pairs. Both architectures are shallow neural models with a linear projection rather than a deep stack of nonlinear hidden layers. Training maintains separate input and output embedding tables, each a matrix containing vectors associated with vocabulary items. The input vectors are typically retained as the word representations. (arxiv.org)

For a central word ww and context word cc, the full softmax formulation is

P(c∣w)=exp⁡(uc⊤vw)∑x∈Vexp⁡(ux⊤vw),P(c\mid w)= \frac{\exp(u_c^\top v_w)} {\sum_{x\in V}\exp(u_x^\top v_w)},

where VV is the vocabulary, vwv_w an input vector, and ucu_c an output vector. Evaluating the denominator over a large vocabulary motivates alternative training objectives. (arxiv.org)

Efficient training objectives

Hierarchical softmax organizes vocabulary items as leaves of a binary tree. Predicting a word involves decisions along the path from the root to its leaf, rather than evaluating every vocabulary item. Word2vec uses a frequency-based Huffman tree, giving common words shorter paths. This reduces the number of output computations required per training example. (arxiv.org)

Negative sampling instead trains a binary discrimination objective. An observed word-context pair is treated as positive, while sampled pairs provide negative examples. For one positive pair, the maximized objective is

log⁡σ(uc⊤vw)+∑i=1klog⁡σ(−uni⊤vw),\log \sigma(u_c^\top v_w) +\sum_{i=1}^{k}\log \sigma(-u_{n_i}^\top v_w),

where σ\sigma is the logistic function, kk is the number of negatives, and nin_i are sampled vocabulary items. Unlike full softmax, this objective does not directly learn a normalized conditional distribution over all words. (arxiv.org)

The original method samples negatives from a distribution proportional to word frequency raised to the 3/43/4 power. It also probabilistically discards some very frequent word occurrences, reducing computation and changing the contexts used for training. Stochastic gradient descent updates the parameters to optimize the chosen objective. Phrase detection can join frequent multiword expressions into single tokens before vector learning. (arxiv.org)

Training choices and implementation

The training corpus must be converted into token sequences through tokenization. Vocabulary thresholds determine which words receive vectors; discarded words and unseen vocabulary items have no learned representation in standard Word2vec. Important hyperparameters include embedding dimension, context-window size, minimum word frequency, negative-sample count, subsampling threshold, number of training passes, and learning rate. These choices affect memory requirements, computation, and the resulting representations. Implementations such as Gensim expose both architectures and the principal training options. (radimrehurek.com)

Geometry and interpretation

Word vectors are often compared using cosine similarity, which measures the angle between vectors rather than their absolute magnitudes. Word2vec research also evaluated analogy completion using vector offsets: a relation between two words can sometimes resemble the offset between another pair. Such regularities are empirical properties of particular embeddings, not guarantees that arbitrary semantic relationships can be solved through vector arithmetic. (radimrehurek.com)

In 2014, Omer Levy and Yoav Goldberg showed that skip-gram with negative sampling has an interpretation as implicit matrix factorization. Under their assumptions, including unigram negative sampling and sufficiently flexible vectors, optimal word-context dot products correspond to pointwise mutual information shifted by −log⁡k-\log k. Low-dimensional embeddings approximate these associations through shared parameters; changing the negative-sampling distribution changes the association being modeled. (papers.neurips.cc)

Limitations and contextual representations

Standard Word2vec learns a static representation for each vocabulary item. Consequently, different senses of the same word share one vector: “bank” has the same representation in financial and riverside contexts. BERT and other contextual models instead produce representations that depend on surrounding text. This distinction separates Word2vec’s stored word-level vectors from context-sensitive representations computed for individual occurrences. (github.com)