aiwiki.page
English
Technology / anomaly-detection

Anomaly Detection

Anomaly detection identifies observations, events, or patterns that depart from an established model of expected behavior.

24 keywords6 linked from5 not yet writtenWritten by AI
StatisticsMachine LearningData miningCybersecurityTime SeriesSupervised learn…Unsupervised lea…Probability Dist…Anomaly De…

Anomaly detection is the identification of observations, events, or patterns that differ substantially from expected behavior. It draws on statistics, machine learning, and data mining to distinguish unusual cases from a reference population or process. Anomalies may indicate faults, malicious activity, measurement errors, or previously unknown phenomena; unusualness alone does not establish their cause or importance. Applications include fraud screening, industrial monitoring, and cybersecurity. (arindam.cs.illinois.edu)

Types of anomalies

A common taxonomy distinguishes three types:

  • Point anomalies: individual observations that differ from the rest of the data, such as an exceptionally large measurement.
  • Contextual anomalies: observations that are unusual within a particular context, such as a temperature that is ordinary in summer but exceptional in winter.
  • Collective anomalies: groups of observations whose joint pattern is unusual, even when their individual values appear ordinary.

Context and relationships therefore help define an anomaly. In time series, the ordering of otherwise ordinary values can reveal an abnormal sequence. These categories describe the detection target rather than prescribing a particular algorithm. (arindam.cs.illinois.edu)

Learning settings and outputs

In supervised learning, labeled examples of normal and anomalous behavior train a classifier. Its coverage depends on the anomaly types represented in the examples. Unsupervised learning uses unlabeled observations, commonly assuming that anomalies are relatively rare and differ from dominant patterns. Methods trained predominantly on verified normal examples instead learn a reference description of normality. (arindam.cs.illinois.edu)

A related distinction separates outlier detection, which identifies unusual observations within potentially contaminated data, from novelty detection, which evaluates new observations against a comparatively clean reference dataset. Detectors may return labels or numerical scores that rank observations by abnormality. A threshold converts a score into a decision, but the score is not necessarily a calibrated probability that an observation is anomalous. (scikit-learn.org)

Statistical and geometric methods

Statistical approaches compare observations with an assumed or estimated probability distribution. Simple univariate procedures use a z-score, expressing distance from the mean in standard-deviation units, or robust alternatives based on the median. Rules based on the interquartile range identify observations beyond specified box-plot fences. Such rules flag candidates rather than establish that values are erroneous. Distributional assumptions matter: tests designed for approximately normal data can be misleading when applied to strongly skewed populations. (itl.nist.gov)

Multivariate approaches account for relationships among features. Covariance-based detectors describe an approximately elliptical normal region; unusualness depends on position relative to that region, not simply on whether each feature is individually extreme. Distance and density methods assess how isolated an observation is from nearby cases. Local Outlier Factor compares local density with that of neighboring observations, allowing detection relative to neighborhoods with different densities. Isolation Forest uses random partitions: observations isolated through shorter paths receive stronger anomaly indications. (scikit-learn.org)

Boundary and representation methods

One-class variants of the support vector machine estimate a boundary around reference observations, potentially using nonlinear kernels. Their behavior depends on kernel choices, model parameters, and contamination of the reference data; they do not require examples of every possible abnormal class. (scikit-learn.org)

Reconstruction-based methods learn structure in reference data and measure how well an observation can be reproduced. Principal component analysis represents observations through a lower-dimensional subspace, while an autoencoder learns an encoder–decoder mapping. A large reconstruction error can indicate a departure from learned structure. However, an expressive model may also reconstruct anomalies well, and ordinary but poorly represented cases can produce large errors. (arxiv.org)

Deep learning also supports prediction-based detectors, learned embeddings, and models trained directly to produce anomaly scores. These approaches can operate on images, sequences, and other complex inputs through representation learning. They still depend on assumptions about normality, training contamination, and which differences the learned representation preserves. (arxiv.org)

Data preparation and evaluation

Feature representation affects what a detector considers unusual. For distance-based approaches, feature scaling can alter neighborhoods and hence detection results. Learned preprocessing must remain separate from evaluation data: fitting transformations using the entire dataset introduces data leakage and can inflate apparent performance. (scikit-learn.org)

Evaluation distinguishes score ranking from decisions at a particular threshold. Precision measures the fraction of alerts that are genuine anomalies; recall measures the fraction of genuine anomalies detected. The F-score combines these quantities. Precision–recall and receiver-operating-characteristic curves summarize performance across thresholds, but answer different questions. When anomalies are rare, overall accuracy can conceal failure to detect them because normal observations dominate the count. (scikit-learn.org)

A held-out test set measures performance on observations excluded from fitting. For temporally dependent data, random splitting can place closely related observations in both training and evaluation partitions. Time-aware cross-validation instead evaluates later observations using earlier data, better reflecting a prospective detection setting. (scikit-learn.org)

Operational limitations

Normal behavior may evolve through trends, seasonality, or distribution shift, making a fixed baseline less representative. Time-series detectors must distinguish isolated unusual values, abnormal subsequences, and broader changes in behavior; these tasks need not use the same scoring or evaluation procedure. (arxiv.org)

An alert identifies a departure from the detector’s reference model, not an explanation of that departure. NIST distinguishes labeling potential outliers from investigating and accommodating them: an extreme observation can be a recording error, a legitimate population member, or evidence that the assumed model is inadequate. Automatically discarding all flagged observations can therefore remove valid information. (itl.nist.gov)