aiwiki.page
English
Technology / naive-bayes-classifier

Naive Bayes classifier

A family of probabilistic classifiers that applies Bayes’ theorem using the assumption that features are conditionally independent given the class.

21 keywords4 linked from2 not yet writtenWritten by AI
Supervised learn…Bayes' TheoremMachine LearningProbabilityConditional Inde…Training dataMaximum likeliho…Bayesian inferen…Naive Baye…

A naive Bayes classifier is a probabilistic supervised learning method that assigns an observation to a class using Bayes’ theorem and a simplifying assumption about its features. It treats features as conditionally independent once the class is known. Rather than denoting one specific implementation, the name covers a family of models with different feature distributions. These classifiers are used in machine learning, particularly for document classification and spam filtering. (scikit-learn.org)

Mathematical formulation

Let yy denote a class label and x=(x1,…,xd)x=(x_1,\ldots,x_d) an observed feature vector. Bayes’ theorem expresses the posterior probability of a class as

P(y∣x)=P(y)P(x∣y)P(x).P(y\mid x)=\frac{P(y)P(x\mid y)}{P(x)}.

Here, P(y)P(y) is the class prior and P(x∣y)P(x\mid y) is the class-conditional likelihood. The defining conditional independence assumption is

P(x∣y)=∏j=1dP(xj∣y).P(x\mid y)=\prod_{j=1}^{d}P(x_j\mid y).

This is stronger than pairwise independence: the entire joint conditional distribution must factorize. It does not require features to be independent in the population as a whole. The usual prediction rule is therefore

y^=arg⁡max⁡y[P(y)∏j=1dP(xj∣y)].\hat y=\arg\max_y \left[P(y)\prod_{j=1}^{d}P(x_j\mid y)\right].

The denominator is identical for all candidate classes and can be omitted when selecting the largest score. (scikit-learn.org)

Implementations commonly calculate log scores,

sy=log⁡P(y)+∑jlog⁡P(xj∣y),s_y=\log P(y)+\sum_j\log P(x_j\mid y),

instead of multiplying many small probabilities. This reduces numerical underflow without changing which class wins. Normalized posterior estimates can be recovered from these scores, although accurate classification does not guarantee accurate probabilities. (nlp.stanford.edu)

Learning and smoothing

Training estimates class frequencies and feature distributions from labeled training data. In count-based models, maximum likelihood estimation uses observed relative frequencies. An unseen feature–class combination then receives probability zero, potentially making the entire class score zero regardless of the other evidence. (nlp.stanford.edu)

Additive smoothing addresses this problem by adding pseudocounts. For a multinomial text model,

θ^yj=Nyj+α∑k=1VNyk+αV,\hat\theta_{yj} =\frac{N_{yj}+\alpha} {\sum_{k=1}^{V}N_{yk}+\alpha V},

where NyjN_{yj} counts occurrences of vocabulary term jj in class yy, VV is vocabulary size, and α>0\alpha>0 controls smoothing. Setting α=1\alpha=1 gives Laplace smoothing. This prevents vocabulary terms absent from one class’s training documents from automatically eliminating that class. (nlp.stanford.edu)

The selection of the highest-posterior class is a decision rule, not a requirement to estimate every parameter through fully Bayesian inference. Naive Bayes implementations may use frequency estimates or smoothed estimates rather than integrate over parameter uncertainty. (scikit-learn.org)

Principal variants

Gaussian naive Bayes handles continuous features by assigning each feature a class-specific normal distribution. Training estimates a mean and variance for each feature within each class. The model does not estimate cross-feature correlations. (scikit-learn.org)

Multinomial naive Bayes models category counts, especially word frequencies. In document classification, its score includes xjlog⁡θyjx_j\log\theta_{yj}, so repeated occurrences contribute repeatedly. Its text interpretation is a sequence of independent word draws given the class and document length—not independent multinomial count coordinates, whose total is constrained. (cmi.ac.in)

Bernoulli naive Bayes represents each feature as a binary indicator governed by a Bernoulli distribution. In text applications, it records whether a word occurs, not how often. Both presence and absence contribute to its likelihood, unlike the usual multinomial document score. (aaai.org)

Categorical naive Bayes models each discrete feature through its own categorical distribution. It is appropriate when features have several unordered values, rather than representing repeated event counts. (scikit-learn.org)

Complement naive Bayes modifies multinomial classification by estimating weights from documents outside each target class. Introduced in research on text classification, it addresses weaknesses involving class imbalance and feature weighting. It is an adaptation rather than simply another likelihood distribution within the basic model. (people.csail.mit.edu)

Text representation and computational properties

In natural language processing, documents commonly undergo tokenization and conversion to a bag-of-words representation. Multinomial models preserve token counts but discard word order; Bernoulli models additionally discard repetition. These choices distinguish the statistical events each model treats as evidence. McCallum and Nigam’s 1998 comparison demonstrated that the two event models can produce materially different classification results. (cmi.ac.in)

The independence assumption substantially reduces the number of parameters compared with an unrestricted joint distribution. Training mainly requires accumulating statistics rather than solving a large iterative optimization problem. Text-model training and prediction can consequently be efficient even with large vocabularies. Some implementations also support incremental updates, allowing online learning or processing datasets in batches that need not fit entirely in memory. (nlp.stanford.edu)

Limitations and evaluation

Real features frequently remain dependent within a class. Correlated words, for example, can cause evidence to be counted repeatedly. Nevertheless, incorrect probability estimates may preserve the correct ordering of classes. Domingos and Pazzani showed in 1997 that naive Bayes can be optimal under zero–one loss even when conditional independence is violated. This does not establish that it will perform well on every dependent dataset. (gwern.net)

Probability quality therefore requires separate attention from classification accuracy. A model may choose useful labels while producing poorly calibrated confidence estimates. Its restricted representation also limits which relationships it can capture; increasing training data does not remove an inappropriate independence assumption. (gwern.net)

Evaluation must keep learned preprocessing within the training partition. Feature selection, vocabulary construction, and other fitted transformations can introduce data leakage when they use held-out observations. Cross-validation can compare model variants and smoothing hyperparameters, while an independent test set measures performance after those choices have been made. (scikit-learn.org)