Anomaly Detection
1. Overview
A. Definition and Background
Anomaly Detection is a data analysis technique that identifies observations (anomalies/outliers) that deviate significantly from the distribution and patterns formed by normal data. It rests on the assumption of imbalance and rarity—that most data is normal and anomalies are rare—learning the boundary of the normal and detecting whatever falls outside it.
Traditional rule-based monitoring relied on human-defined conditions such as "raise an alert when a threshold is exceeded." However, as digital transformation led servers, networks, IoT sensors, and financial transactions to produce tens of thousands of logs and time series per second, having humans enumerate every normal pattern as a rule hit its limits. The normal state itself changes with time, load, and seasonality, and attackers and fraudsters evolve to bypass known rules.
Anomaly detection reframes this problem: instead of enumerating "what is abnormal," it learns "what is normal" from data and quantifies the degree of deviation from that normal model (the anomaly score). Being able to capture even unknown (zero-day) threats, failures, and fraud is the decisive difference from rule-based approaches. That it is mostly approached as unsupervised or semi-supervised learning in reality, where labels are scarce, is another characteristic that distinguishes it from ordinary classification problems.
B. Necessity and Characteristics
Anomaly detection is needed because the loss caused by a single anomalous event is asymmetrically large. Catching one fraudulent card transaction, one subtle vibration anomaly in a manufacturing machine, or one abnormal traffic flow at the start of an intrusion prevents a major incident, but missing it chains into financial, safety, and reputational damage. Conversely, mistaking normal for abnormal (a False Positive) drives operators into Alert Fatigue, leading them to ignore real alerts. Thus anomaly detection takes the balance between False Negatives and False Positives as a core design goal.
Its characteristics can be summarized in three points. First, extreme class imbalance: anomalies are usually 0.1–1%, so Accuracy is meaningless and metrics such as precision, recall, and PR-AUC are needed. Second, the scarcity of ground-truth labels, making supervised learning hard to apply directly. Third, context dependence: the same value can be normal or abnormal depending on time and surrounding conditions (e.g., a large payment at 3 a.m.).
Because of these traits, anomaly detection must be framed differently from ordinary prediction and classification tasks from the outset. Whereas classification asks "is this data A or B," anomaly detection is closer to an open-set problem asking "does this data belong to the learned normal world." Since it must capture even new, undefined anomalies, how precisely it estimates the boundary of the normal governs performance, which in turn reduces to the quality of data representativeness and feature design.
2. Types of Anomalies and the Detection Process
To design anomaly detection, one must first define the nature of the target anomaly. Anomalies fall broadly into three kinds. A Point Anomaly is when an individual observation deviates alone from the overall distribution—for instance, a sudden payment of 5 million won amid a history of payments in the 50,000-won range. A Contextual Anomaly is when the value itself is within a normal range but abnormal in a specific context—electricity usage normal in summer is treated as anomalous if it appears at midnight in winter. A Collective Anomaly is when individual values are normal but the entire consecutive pattern is abnormal—the classic case being an ECG in which individual waveforms are within normal range but the rhythm pattern over a segment indicates arrhythmia.
This distinction matters because the appropriate techniques and feature design differ by type. Point anomalies are caught by the distribution of the values themselves, but contextual anomalies must model context variables such as time and location together, and collective anomalies require representation learning at the sequence/window level. Misdefining the type means even the best algorithm will miss them.
The overall detection pipeline forms a single circular structure from data collection to response. The structure diagram below shows a closed loop that learns a normal model, computes anomaly scores, decides via a threshold, and then updates the model with operator feedback.
flowchart LR
A["Data collection(logs/time series/transactions)"] --> B["Preprocessing·normalization·missing handling"]
B --> C["Feature extraction(Feature Engineering)"]
C --> D["Normal model training"]
D --> E["Anomaly Score computation"]
E --> F["Threshold·decision"]
F --> G["Alerting·visualization·response"]
G --> H["Operator feedback·labeling"]
H --> D
| Anomaly type | Definition | Example | Key design point |
|---|---|---|---|
| Point anomaly | A single value deviates from the distribution | High-value single payment | Value distribution·density model |
| Contextual anomaly | Abnormal in context | Heavy traffic at midnight | Context variables such as time·location |
| Collective anomaly | An abnormal pattern·sequence | Arrhythmia waveform, gradually rising latency | Window·sequence representation |
In the preprocessing stage, normalization (scaling) and handling of missing values and noise are especially important. Distance- and density-based techniques are sensitive to scale, so without standardization, variables with large scales dominate the distance computation. In the feature extraction stage, for time series, moving averages, rates of change, and periodic components (FFT), and for logs, aggregate statistics per session or entity, determine detection power.
3. Classification of Anomaly Detection Techniques
Anomaly detection techniques can be divided—by how they represent the normal—into statistical, distance/density-based, tree-based, and deep-learning-based. The detailed architecture below shows a practical structure that combines multiple detectors in an ensemble within one system and integrates their scores to decide.
flowchart TB
subgraph Input["Input"]
X["Normalized feature vector"]
end
subgraph Detectors["Detector ensemble"]
S1["Statistical(Z-score·IQR)"]
S2["Distance·density(kNN·LOF)"]
S3["Tree(Isolation Forest)"]
S4["Deep learning(Autoencoder·LSTM)"]
end
X --> S1
X --> S2
X --> S3
X --> S4
S1 --> AGG["Score integration·normalization"]
S2 --> AGG
S3 --> AGG
S4 --> AGG
AGG --> T["Threshold·rank-based decision"]
T --> O["Anomaly list·explanation(XAI)"]
A. Statistical Techniques
Statistical techniques assume that data follows a particular distribution (mainly Gaussian) and treat regions that are probabilistically sparse in that distribution as anomalies. The simplest, the Z-score, computes how many standard deviations a point is from the mean and typically flags an absolute value above 3 (about the top/bottom 0.13%) as anomalous. When the distribution is skewed and the mean/standard deviation are distorted by extremes, the IQR (Interquartile Range) method using the median and quartiles is robust. IQR defines below Q1-1.5×IQR or above Q3+1.5×IQR as anomalous.
The strength of statistical techniques is that computation is light and the statistical meaning of the threshold is clear, easing regulatory and audit responses. In practice, the first-stage filter of a financial fraud detection system (FDS) or the control chart (±3σ) of a manufacturing process are representative. However, they are vulnerable to multivariate/nonlinear relationships or data where the distribution assumption breaks, so for complex data such as high-dimensional logs or images they must be used alongside the techniques described below.
A caution here is that one must always verify whether the distribution assumption holds. For example, network packet sizes or response times are mostly long-tailed (with a long right tail) distributions, so applying a Z-score that presumes a Gaussian will over-flag the normal tail en masse. In such cases, apply it after a log transform or switch to a quantile-based technique.
B. Distance/Density-Based Techniques
Distance/density-based techniques formalize the intuition that "outliers lie far from normal clusters or in low-density regions." The kNN-based approach takes the distance to the k-th nearest neighbor as the anomaly score, treating large distances as isolated anomalies. LOF (Local Outlier Factor) goes a step further, computing the ratio of one's own density to that of neighbors, thereby catching points that are locally—not globally—low in density. On real data with non-uniform density (e.g., where payment-pattern density differs between city and suburb), LOF outperforms global distance techniques.
The strength of this family is that it works from the geometric structure of data alone, with no distribution assumption. On the other hand, its weaknesses are the Curse of Dimensionality, where the notion of distance becomes meaningless in high dimensions, and a computational load reaching O(n²) when computing all pairwise distances. Therefore, at large scales of millions of records or hundreds of dimensions, it is combined with dimensionality reduction (PCA·autoencoder) or approximate-nearest-neighbor (ANN) indexes.
As a practical example, in a telecom's roaming-fraud detection, vectorizing each subscriber's usage pattern and applying LOF can catch usage spikes that are locally different from the norm. Here the choice of k (number of neighbors) is sensitive—too small is vulnerable to noise, too large loses locality—so it is tuned with validation data.
C. Tree (Isolation)-Based Techniques
Isolation Forest exploits the property that outliers are "easily isolated with few splits." When binary trees are built repeatedly by randomly choosing features and split points, outliers are separated alone at shallower depths than normal points, so their average path length is short. The normalized path length becomes the anomaly score. The key shift in thinking is that it directly isolates anomalies without needing to model the dense region of the normal.
The practical appeal of Isolation Forest is its near-linear O(n) scalability, freedom from distribution assumptions, and relative robustness in high dimensions. It works well with just sampling (sub-sampling), so it is widely used for first-pass screening of large-scale log and transaction data. However, since it makes only axis-parallel splits, it struggles with complex boundaries oblique to the axes, in which case it is complemented by the Extended Isolation Forest or deep-learning techniques.
D. Deep-Learning/Machine-Learning Techniques
For complex nonlinear, high-dimensional, and time-series data, representation-learning-based techniques are powerful. An Autoencoder is trained only on normal data to compress input into a low dimension and then reconstruct it, after which inputs whose Reconstruction Error is large because they reconstruct poorly are judged anomalous. Having learned only the normal, it cannot reconstruct never-before-seen anomaly patterns well—that is the principle. For time series, LSTM/Transformer-based prediction/reconstruction models predict the next value and flag segments with large prediction error as anomalies. For images and industrial inspection, methods that learn the distribution of normal images and then detect deviation are used.
This family has the advantage of capturing complex patterns (including collective anomalies) without humans designing the raw features. In return, it demands much training data and computation, and it is hard to explain why something is anomalous, so combination with XAI (explainable AI) techniques is needed. The semi-supervised One-Class SVM learns the minimal boundary enclosing normal data and treats outside the boundary as anomalous, effective for small-to-medium normal sets.
There are also applications of generative models. GAN-based anomaly detection (the AnoGAN family) learns the generative distribution of normal data, then finds the latent vector that best reproduces the input and takes that reproduction difference as the anomaly score; recently, research on using diffusion models to reconstruct the normal and catch deviation is also active. However, the more advanced such a model, the greater the burden of training stability, interpretability, and computational cost, so avoiding excessive complexity in view of the problem's difficulty and operational constraints is a practical principle. That is, a hierarchical strategy that handles most cases with statistical/isolation techniques and selectively deploys deep learning for the complex, time-series anomalies they miss is cost-effective.
4. Comparison of Techniques and Application Cases
Technique choice is decided not by "which has the highest performance" but at the intersection of constraints: data scale, dimensionality, presence of labels, explanation requirements, and real-time needs. Statistical techniques are light and easy to explain but low in expressive power; distance/density techniques reflect structure well but are weak on high-dimensional/large-scale data. Isolation Forest is strong in scalability and assumption-freedom but limited on complex boundaries, and deep learning has the highest expressive power but large demands on resources, explainability, and data. This difference arises because each technique represents "normal" in a different way (distribution, distance, isolation, reconstruction).
| Technique | Labels needed | High-dim | Scalability | Explainability | Representative use |
|---|---|---|---|---|---|
| Statistical(Z·IQR) | Not needed | Low | Very high | High | Control chart, first-stage filter |
| Distance·density(LOF) | Not needed | Low | Low(O(n²)) | Medium | Local anomalies, small/medium scale |
| Isolation Forest | Not needed | Medium | High(O(n)) | Medium | Large-scale screening |
| Autoencoder/LSTM | Semi-supervised | High | Medium | Low | Time series·images·complex patterns |
| One-Class SVM | Semi-supervised | Medium | Low | Medium | Small/medium normal sets |
The first case is financial fraud detection (FDS). Domestic card companies typically use a multi-stage structure: a rule-based first-stage filter screens out obvious anomalies to reduce false positives, followed by a machine-learning (Isolation Forest·autoencoder) second-stage model to catch subtle fraud patterns. With a fraud rate hovering around 0.1%, extremely imbalanced, the crux is tuning the threshold to raise recall while suppressing the blocking of normal transactions (customer inconvenience) caused by false positives.
The second case is predictive maintenance (PdM) of manufacturing equipment. By learning normal operating patterns from vibration, temperature, and current sensor time series with an autoencoder or LSTM, a surge in reconstruction error can capture early signs of failure such as bearing wear days to weeks in advance. In practice, early detection of equipment anomalies allows a switch to planned maintenance, meaningfully reducing unplanned downtime.
The third case is IT operations and security monitoring (AIOps·UEBA). Applying anomaly detection to server metrics, logs, and user behavior (UEBA) captures traffic spikes, abnormal logins, and insider threats. For example, if an account that normally connects only during business hours from a specific IP attempts a mass download from an overseas IP at midnight, it is caught as a contextual anomaly and automatically responded to via SIEM/SOAR.
The common lesson running through these three cases is that rather than seeking a single all-purpose algorithm, it is more robust to design features with domain knowledge and combine multiple detectors in a hierarchy/ensemble. The FDS multi-stage filter, PdM's combination of per-sensor models, and monitoring's rules+ML in parallel all follow the same principle. Moreover, all three cases show the common trait that, more than detection itself, the closed loop of operator review, label accumulation, and model updating after detection governs long-term performance.
5. Deep Dive: Recent Trends and Evaluation Strategy
Recent trends point in three directions. First, use of deep learning and foundation models: industrial-inspection models learning the distribution of normal images, Transformers for time series, and graph-neural-network (GNN)-based anomaly detection for financial and social networks are spreading. In particular, approaches that leverage pretrained representations to boost detection performance even with small data draw attention. Second, combination with explainability (XAI): in regulated industries (finance, healthcare), one must present "why it is anomalous," so the grounds for detection are provided together via SHAP, reconstruction-error contribution analysis, and the like. Third, real-time streaming and Concept Drift response: since normal patterns change over time, online learning and a model-retraining pipeline (MLOps) have become essential.
The design of evaluation metrics decides success or failure in anomaly detection. Because classes are extremely imbalanced, accuracy misleads (predicting all-normal still yields 99%). Instead, use precision, recall, F1, and PR-AUC, prioritizing recall in domains where nothing may be missed (safety·security) and precision in domains where alert fatigue is severe. In time series, since anomalies exist as "segments," look at segment-level evaluation or Detection Delay together rather than point-level metrics. Also, the threshold is set not as a fixed value but from a cost-sensitive perspective reflecting operational costs (cost of handling false positives versus loss from false negatives).
6. Considerations and Implications
From a professional engineer's perspective, adopting anomaly detection is a strategic decision spanning data, operations, and governance beyond algorithm choice. First, a label and data-quality strategy. Anomaly labels are scarce and delayed (months until fraud is confirmed), so a realistic roadmap is to start unsupervised, accumulate labels through operator feedback, and gradually advance to semi-supervised and active learning. Initial data cleansing must proceed in parallel so that anomalies mixed into normal data do not contaminate learning.
Second, cost-based management of the false-positive/false-negative trade-off. Since a single threshold makes the two errors inversely related, define a per-domain loss function to set the threshold and absorb false positives with a multi-stage structure (rules → ML → human review). To reduce alert fatigue, grading alerts by severity and correlating/aggregating similar alerts is the crux of operational success.
Third, explainability and regulatory/audit response. In finance, healthcare, and personal-data domains, presenting the grounds for automated decisions and an appeal process is required. Therefore, rather than a high-performing black box, choose a model that can present grounds or an XAI-combined structure, and keep decision histories, model versions, and data lineage auditable.
Fourth, operational continuity and adjacent technologies. Completing the "detection-to-response" closed loop through a retraining pipeline that responds to concept drift (MLOps·model drift monitoring), real-time streaming processing (Kafka·Flink), and linkage with automated response after detection (SOAR) is what yields real value. Also, to prepare for Adversarial Evasion, it is advisable to ensemble diverse detectors and periodically test detection-bypass scenarios. Looking ahead, foundation models and self-supervised learning will ease the label-shortage problem, and anomaly detection will integrate with Observability and Data Observability to establish itself as a standard component of system reliability.
References
- Chandola, Banerjee, Kumar, "Anomaly Detection: A Survey", ACM Computing Surveys, https://dl.acm.org/doi/10.1145/1541880.1541882
- scikit-learn, "Novelty and Outlier Detection", https://scikit-learn.org/stable/modules/outlier_detection.html
- Liu, Ting, Zhou, "Isolation Forest", https://ieeexplore.ieee.org/document/4781136
In one line: Anomaly detection learns "normal" from data to detect rare, abnormal observations outside its boundary, and it delivers practical value only when point/contextual/collective anomaly types and statistical/distance/isolation/deep-learning techniques are combined to fit data constraints and cost, while the false-positive/false-negative balance, explainability, and an operational closed loop are designed together.