← Back to list
AI & Data
#SOM#자기조직화지도#군집분석#비지도학습#Kohonen#134회
Last updated · 2026-09-27

SOM (Self Organizing Map)

1. Overview

A. Definition

A SOM (Self Organizing Map) is an unsupervised learning neural network proposed by Finland's Teuvo Kohonen (1982) that clusters, reduces dimensionality, and visualizes high-dimensional input data by mapping it onto a low-dimensional (usually 2D) grid while preserving topology. It is also called a Kohonen Map or Kohonen network.

The core idea of SOM comes from the way the human cerebral cortex works. In the brain, a phenomenon called the topographic map is observed, in which similar sensory stimuli (e.g., touch from adjacent fingers) are processed in adjacent neuron regions of the cerebral cortex, and SOM artificially imitates this principle. That is, the network self-organizes so that similar inputs correspond to nearby neurons on the output grid. As a result, the complex neighborhood relationships of high-dimensional data are laid out on a two-dimensional map in a form that people can inspect with their eyes. This "topology preservation" is the decisive feature that distinguishes SOM from simple clustering techniques.

B. Background and Necessity

As the dimensionality of data grows, the 'curse of dimensionality' arises, making it harder for people to grasp structure intuitively. Looking directly at customer, gene, or document data represented by tens to hundreds of variables, one cannot tell which things are similar to each other or what clusters exist. Traditional clustering (e.g., K-means) divides data into a few groups but cannot reveal the mutual relationships among clusters (which cluster is adjacent to which). SOM performs clustering and visualization at the same time while also showing the topological relationships among clusters on a two-dimensional map, thereby satisfying both requirements of exploratory data analysis and visual communication of results.

In addition, linear dimensionality reduction such as PCA projects data onto orthogonal axes of large variance, so it has the limitation of not properly unfolding nonlinear structures that curve into a surface shape. SOM, in which grid neurons place themselves along the data distribution, adds to its necessity by being able to capture nonlinear manifold structures as well. In particular, because the result comes out not as 'a scatter of points' but as 'a grid with interpretable prototype vectors', it shows strength in practical situations where analysis results must be explained and argued to non-expert stakeholders.

C. Characteristics

SOM is an unsupervised method that learns the internal structure of data without ground-truth labels, and it operates through competitive learning, in which neurons compete over an input to select a winner. In particular, unlike ordinary competitive learning that updates only the winner, the essence of SOM is that it updates the winner's neighbor neurons together, which makes adjacent neurons take on similar values and realizes topology preservation. If there were no neighbor update, SOM would amount to mere vector quantization (competitive learning), in which neurons scatter individually to approximate the data, and the grid position would carry no meaning. In other words, the single device of 'updating neighbors together' elevates vector quantization into a topology-preserving map. Also, because it is a shallow (two-layer) structure with no hidden layers, backpropagation is unnecessary and interpretation is relatively easy.

Characteristic Description
Unsupervised learning Learns data structure without ground-truth labels
Topology preservation Maintains adjacency relationships of the input space on the output grid
Competitive learning Selects a winner (BMU) through competition among neurons
Neighbor update Updates the winner and its neighbors together (key differentiator)
Dimensionality reduction & visualization Clusters and visualizes high dimensions on a 2D grid (nonlinear)

2. SOM Structure and Components

flowchart LR
  subgraph IN["Input Layer"]
    X["high-dimensional input vector x = (x1..xn)"]
  end
  subgraph OUT["Competitive/Output Layer (2D grid)"]
    N1(("neuron"))
    N2(("neuron"))
    N3(("BMU winner"))
    N4(("neuron"))
  end
  X -->|"weight vector w (fully connected)"| N1
  X --> N2
  X --> N3
  X --> N4
  N3 -.->|"updated together via neighborhood function"| N2
  N3 -.-> N4
  style N3 fill:#e8f0fe,stroke:#2f6fed

Unlike a general deep neural network with multiple hidden layers, SOM has a simple structure consisting of just two layers: the input layer and the competitive layer (output layer). The input layer is merely a passage that receives high-dimensional input vectors and performs no computation, while the actual learning takes place in the competitive layer. The competitive layer is usually composed of neurons arranged in a rectangular or hexagonal grid, and each neuron holds a weight (reference) vector of the same dimension as the input. That is, if the input is 20-dimensional, every neuron also has a 20-dimensional weight vector, and this vector becomes the 'prototype' pattern that the neuron represents.

Structurally, the input layer and the competitive layer are fully connected, so a single input vector is delivered to all competitive-layer neurons simultaneously. However, the difference is that, rather than propagating a weighted sum to the next layer as in a general neural network, the connections are used only for each neuron to compute the 'similarity (distance)' between its own weight vector and the input. In other words, in SOM the connection weights are not a channel for transmitting signals but the representative pattern the neuron remembers, in itself. Understanding this difference in perspective makes it natural to see why SOM learns even without backpropagation.

When an input arrives, SOM computes the distance (usually Euclidean distance) between the input vector and the weight vector of every neuron, and selects the nearest neuron as the winner, i.e., the BMU (Best Matching Unit). This winner selection corresponds to 'competition'. Then it pulls not only the BMU but also the weights of neighboring neurons around the BMU on the grid toward the input. Here, how strongly each neighbor is pulled is determined by a neighborhood function (usually Gaussian) that decreases with the grid distance from the BMU. Near neighbors are updated strongly and far neighbors weakly, so as learning is repeated, adjacent neurons on the grid come to have similar weights and topology is preserved. A point to note here is that the 'neighborhood' is defined by position on the output grid, not in the input space, and this is the mechanism that projects high-dimensional neighborhood relationships into two dimensions.

Component Description
Input layer Passage that receives high-dimensional input vectors (no computation)
Competitive layer (output layer) 2D grid neurons, each holding a weight vector
Weight vector Reference vector compared with the input (the neuron's prototype pattern)
BMU (Best Matching Unit) The winner neuron most similar to the input
Neighborhood function Updates neurons around the BMU together (Gaussian), radius decreases
Learning rate Magnitude of weight movement, decreases as learning proceeds

The choice of distance metric is also a factor that changes the result. Standard SOM uses Euclidean distance, but cosine distance can be used for text/sparse vectors and Mahalanobis distance for mixed-type variables with different scales; when the metric changes, the definition of 'similarity' changes, and so does the cluster boundary of the map. Therefore, designing a distance metric and variable scaling suited to the domain characteristics together is a prerequisite for obtaining a good map.

The shape and size of the grid are also design factors. If neurons are arranged in a rectangular grid, each neuron's neighbors are defined as 4 (up/down/left/right, or 8 including diagonals); if arranged in a hexagonal grid, there are 6 neighbors, reducing directional bias and making topology representation more natural, which is preferred when visualization quality matters. The grid size, i.e., the total number of neurons, determines the map's resolution: too small and different clusters clump into one neuron and cannot be distinguished, too large and there are many empty neurons with no data and learning cost grows. There is a practice of starting from a level proportional to the square root of the number of data points (e.g., about 5√N of the sample size) and adjusting, but this should be seen not as an absolute rule but as a starting point to tune to the objective and data.

3. Learning Procedure

flowchart TB
  A["initialize weights<br/>(random or PCA-based)"] --> B["present input vector x"]
  B --> C["compute distance to all neurons"]
  C --> D["select the minimum-distance neuron as BMU"]
  D --> E["update BMU and neighbor weights<br/>toward the input"]
  E --> F["shrink learning rate and neighborhood radius"]
  F --> G{"converged or<br/>max iterations reached?"}
  G -->|"no"| B
  G -->|"yes"| H["end learning<br/>(topology-preserving map complete)"]

The core of SOM learning is to "repeatedly select a winner and move the winner and its neighbors a little toward the input". The weight update rule is conceptually of the form new weight = old weight + learning rate × neighbor strength × (input − old weight), pulling the neuron's prototype vector little by little toward the observed input. The neighbor strength is generally given in the Gaussian form exp(−d²/2σ²) with respect to the grid distance d from the BMU and the current neighborhood radius σ, being close to 1 at the BMU itself and converging to 0 as distance grows. This rule is a competitive variant of Hebbian learning, implementing in a formula the self-organizing principle that 'a frequently-winning neuron and its neighbors specialize to represent that input pattern'. The decisively important point here is that as learning proceeds, the learning rate and neighborhood radius are gradually shrunk.

This gradual shrinkage can be understood in two phases. In the initial ordering phase, a wide neighborhood radius and large learning rate lay out the broad framework of the whole map, so randomly initialized weights are globally aligned to the rough shape of the data distribution. In the following fine-tuning (convergence) phase, the neighborhood radius is narrowed almost to the BMU itself and the learning rate is lowered, so each neuron precisely refines the detailed structure of its region of responsibility. If the neighborhood radius is kept narrow from the start, the map twists (topological defect) and topology breaks; conversely, if kept wide to the end, the map smears and cannot distinguish detailed clusters. Therefore the shrinkage schedule of radius and learning rate is a key hyperparameter that governs SOM quality.

This two-phase scheme corresponds to the 'exploration then exploitation' strategy spoken of in optimization in general. The large learning rate and wide radius early on encourage broad exploration so as not to prematurely fall into a local optimum, and the small values later stably converge the found structure. Learning quality is conventionally checked with two quantitative metrics. One is the quantization error, the average distance between each input and its BMU, which indicates how well the map approximates the data; the other is the topographic error, the proportion of inputs whose 1st and 2nd BMUs are adjacent on the grid, which indicates how well topology is preserved. The two metrics often conflict (raising resolution may shake topology), so finding a balance point is the goal of tuning.

Order Description
1 Initialize weight vectors (random or data-based)
2 Present the input vector
3 Compare distance to all neurons → select the minimum-distance neuron as BMU
4 Update BMU and neighbor weights toward the input
5 Shrink learning rate and neighborhood radius, repeat 2-4
6 Terminate on convergence, visualize with U-Matrix etc.

The initialization method also affects convergence speed and quality. If weights start completely at random, the map oscillates greatly in the initial ordering phase and convergence is slow; but if the grid is placed linearly on the plane spanned by the two principal-component (PCA) axes of the data to start, it begins already roughly aligned to the data distribution, so it converges much faster and more stably. This is why practical libraries offer PCA-based initialization as an option.

There are two learning modes: online (sequential) learning, which presents inputs one at a time and updates immediately, and batch learning, which considers all data at once and updates each neuron's weight to the (neighbor-weighted) average of its responsible inputs. The online mode is somewhat sensitive to the order of input presentation but is good for applying to streaming data, while the batch mode has the practical advantage of no order dependence, easy parallelization, and fast convergence on large data. In either mode, the goal of topology preservation and the principle of shrinking the radius and learning rate are identical.

A trained SOM yields not a simple clustering result but a single 'map', and the representative tool for interpreting it is the U-Matrix (Unified Distance Matrix). The U-Matrix expresses the weight distance between adjacent neurons as color (shade): boundaries with large distance appear dark, revealing the 'valleys' between clusters, and regions with small distance appear light, revealing the 'basin' of a single cluster visually. Beyond this, using together a hit map, which shows the number of inputs mapped to each neuron, and a component plane, which colors each neuron by the value distribution of a specific variable, one can even interpret which variable characterizes which cluster. Through such visualizations, the analyst grasps at a glance how many natural clusters exist in the data and how the clusters are adjacent to each other.

4. Differences Between SOM and Other Techniques — Comparison and Cases

SOM and general supervised neural networks such as MLP (multilayer perceptron) differ in their very purpose. Whereas MLP is told the answers and learns classification/regression by backpropagating prediction error, SOM grasps the structure of data through competitive learning (winner-take-all) without answers to cluster and visualize. The absence of error backpropagation and the output of a topology-preserving 2D map are the fundamental differences. The cooperation mode of learning is also contrasting. Supervised learning has all neurons share the error and adjust together, whereas SOM takes a local cooperation in which only the winner and its neighbors are updated and far neurons remain as they are. This locality is the driving force that makes different regions on the map represent different patterns by division of labor. Meanwhile, compared with K-means, another unsupervised clustering, K-means fixes the number of clusters K in advance and merely assigns each point to the nearest centroid without expressing relationships among clusters, whereas SOM places many neurons (corresponding to K-means centroids) on a grid and, through their adjacency, shows even the topology among clusters. In fact, SOM is also interpreted as 'a generalization of K-means with a topological constraint'.

Concrete cases make SOM's practical value clear. First, in customer segmentation, training SOM on tens of purchase/behavior variables gathers similar customer groups into adjacent regions on the map, letting one visually grasp marketing target groups and the transition relationships among them. For example, mapping 100,000 customers represented by 30 variables onto a 20×20 (400 neurons) grid lets one confirm with the eye where 'high-frequency, high-value buyers' and 'churn-risk buyers' are adjacent on a map summarized to 400 dimensions or fewer, and design a transition marketing strategy. Second, in financial distress/credit risk analysis, tens of corporate financial ratios are spread onto a two-dimensional map to use for early warning of which regions distress-signal companies fall into (a well-known research case applied SOM to actual corporate bankruptcy analysis in Finland). Third, in anomaly detection of manufacturing processes/equipment, if a map is built from normal-operation data, when a new observation maps to a position far (large distance from its BMU) from the normal region on the map, it can be caught as an anomaly signal. All three cases exploit SOM's characteristic of providing 'clustering + relationship visualization' at once.

Here it is worth pointing out again why the difference between SOM and K-means is important in practice. Applying K-means to the same customer data yields labels like 'clusters 1-5', but one cannot tell whether cluster 2 is a neighbor of cluster 3 or the opposite, so a marketer finds it hard to design transition paths among clusters. In contrast, SOM places adjacent clusters into adjacent regions on the map, so it visually supports the story that "a customer moves from this state to that state". This 'visibility of relationships' is the point that makes SOM a more persuasive communication tool than simple clustering.

Category SOM K-means Supervised such as MLP
Learning mode Unsupervised Unsupervised Mostly supervised
Learning rule Competitive learning (winner + neighbors) Repeated centroid recomputation Error backpropagation
Purpose Clustering/dimensionality reduction/visualization Cluster partitioning Classification/prediction
Inter-cluster relationship Expressed by topology Cannot express Not applicable
Output Topology-preserving 2D map Cluster labels/centroids Class/continuous value

Meanwhile, SOM does not stay purely unsupervised. Semi-supervised use is possible, in which after learning a small number of labels are assigned to each neuron (cluster) to classify a new input by the label of its neuron, and Kohonen also proposed a supervised variant (LVQ, Learning Vector Quantization) that partially reflects label information in learning. That is, the SOM family forms a spectrum from unsupervised visualization to semi-supervised classification, and is flexibly applied according to the nature of the problem.

5. Deep Dive: Variant Techniques and Recent Trends

SOM is a classic technique established in the 1980s-90s, but it continues through various variants and modern reinterpretations. Basic SOM must fix the grid size (number of neurons) in advance, but the Growing SOM/GHSOM (Growing Hierarchical SOM), in which the grid grows to fit the data, eases this constraint by adding and hierarchizing neurons where needed. GHSOM in particular, when the data has a complex hierarchical structure, automatically reveals major-category to minor-category by having one neuron of an upper map unfold into a lower map, and is used in document classification, web log analysis, and the like. In addition, GTM (Generative Topographic Mapping, Bishop et al.), which reformulates SOM's heuristic nature as a probabilistic model, provides an explicit objective function and probabilistic interpretation, improving convergence and comparability. Beyond these, various extensions exist, such as recursive SOM for structured data like time series/strings, and variants that adjust the distance metric through learning.

It is also worth noting that SOM's learning has no single loss function that is explicitly minimized. Basic SOM is closer to a heuristic in which organization emerges from the repetition of local rules of competition and cooperation, so quantitative comparison among different initializations/parameters is tricky. This is why GTM reformulated this SOM as a generative probabilistic model, a Gaussian mixture over latent variables, thanks to which one can compare models by the clear measure of likelihood and handle missing values and uncertainty in a principled way. In practice, one chooses a technique considering this trade-off between theoretical rigor and implementation simplicity/interpretability.

Recently, as nonlinear dimensionality-reduction/visualization techniques such as t-SNE and UMAP have come into wide use, SOM's visualization role has been partly replaced, but SOM still has unique strengths. Whereas t-SNE/UMAP merely scatter each data point into two dimensions, SOM learns a stable coordinate system (codebook) of grid neurons, so it can immediately map new data onto the existing map (online classification), and each neuron holds a representative prototype vector, making results easy to interpret. For example, t-SNE in principle must recompute the whole thing when new data arrives, but a trained SOM only needs to find the BMU of the new observation, so it suits real-time streaming environments. For this reason it continues to be used practically in equipment-state monitoring, quality control, and vector-quantization (VQ)-based signal/image compression in industrial settings. Research toward combining with deep learning is also underway, such as attempts to re-visualize deep learning's latent representations (embeddings) with SOM so that people can interpret the structure of the feature space a deep model has learned.

6. Considerations and Implications

  1. Clarifying the role — it is an exploration/explanation tool, not a prediction model. From a professional engineer's perspective, SOM is appropriately positioned not as a final model that produces accurate predictions but as an exploratory data analysis (EDA)/explanation tool that understands data structure and visually communicates with stakeholders ahead of full-scale modeling. It should be combined complementarily with K-means, PCA, t-SNE, and the like.

  2. Managing hyperparameter sensitivity is essential. Because results depend greatly on grid size, initial learning rate, neighborhood radius and its shrinkage schedule, and the number of learning iterations, map quality must be evaluated with quantitative metrics such as topographic error and quantization error and tuned iteratively. For reproducibility, a system that records and manages the initialization method and parameters is needed.

  3. Data preprocessing governs the result. Since SOM is distance-based, a large scale difference among variables lets a particular variable dominate the distance computation. Standardization (normalization) and handling of missing values/outliers must precede it, and categorical variables need appropriate encoding.

  4. Reviewing scalability and alternative techniques. When the data scale is very large or ultra-high-dimensional, there are limits in learning cost and interpretation, so a comparative review with a mini-batch/parallel implementation or alternatives such as GTM/UMAP is needed. Choose to fit the objective: UMAP is advantageous if the purpose is pure visualization, SOM if online mapping/prototype interpretation matters.

  5. Responsibility for interpreting and validating results. Because the cluster boundaries a SOM map shows are a product of the data and parameters, one must validate with domain experts whether the derived clusters are actually meaningful for the business and guard against over-interpretation.

  6. Model governance/reproducibility perspective. When incorporating SOM into an operational system such as anomaly detection, a management system is needed that records the data/parameters/version used in learning and periodically retrains the map according to data distribution change (drift). To leverage its strength as an explanation tool, the interpretive basis of the map (component planes, U-Matrix) should be documented together to secure the traceability of decisions.

In sum, SOM is not a technique that competes on flashy prediction performance, but it still holds a hard-to-replace position for the purpose of 'making high-dimensional data understandable to people'. As a professional engineer, one is required to have the capability to judge at which stage and for what purpose to place SOM within an analysis pipeline where deep learning, traditional statistics, and the latest visualization techniques coexist, and to clearly communicate the reliability and limits of its results.

References


In one line: SOM is an unsupervised neural network based on competitive learning that maps high-dimensional data onto a 2D grid while preserving topology (clustering/visualization); by updating the winner (BMU) and its neighbors together and converging by shrinking the learning rate and radius, it differs in purpose and learning rule from backpropagation-based supervised learning and from K-means, which cannot see relationships, and it is used in practice as an exploratory-analysis/explanation tool.