Last updated: 2026-10-07
Self-Organising Maps
Every model up to this page has been supervised — trained against known labels. A Self-Organising Map (SOM), introduced by Kohonen in 1982, is this section's first genuinely unsupervised architecture: it is handed unlabelled data and builds its own low-dimensional map of it, with no target output to train against at all1.
The Grid and the Best-Matching Unit FoundationalKnowledge that endures for decades — core principles
A SOM is a grid of units — commonly a 2D rectangle, however many dimensions the actual input data has — where every unit holds its own prototype vector in the input's original space. Training repeats a simple cycle for each input example: find the best-matching unit (the grid unit whose prototype vector is closest to this example), then update that unit's prototype slightly to be even closer to the example, and update its grid neighbours too, by a smaller amount that fades with distance across the grid.
This is competitive learning: grid units compete for the right to represent each input, and the winner (plus its shrinking neighbourhood) gets to move toward it — "winner-take-most," since nearby units still get a smaller share of the update, not "winner-take-all." Both the learning rate and the neighbourhood radius are shrunk gradually over training, so the map makes large, coarse adjustments early on and small, fine ones later.winner-take-most, not winner-take-all
Topology Preservation as a Goal, Not a Guarantee FoundationalKnowledge that endures for decades — core principles
Because neighbouring grid units are pulled toward similar inputs, units that end up close together on the grid tend to represent similar regions of the original input space — the map develops a topology-preserving layout, where nearby on the grid roughly means nearby in the data. This is the property that makes a trained SOM useful as a 2D visualisation of high-dimensional data: points that were close in the original space tend to land near each other on the map. It is an emergent tendency the training procedure encourages, not a mathematical guarantee — a map can still develop twists or folds where the 2D grid can't fully respect every distance relationship present in genuinely high-dimensional data, the same way any 2D projection of a higher-dimensional shape must distort something.
Clustering, Dimensionality Reduction, Visualisation — and Not Classification FoundationalKnowledge that endures for decades — core principles
A SOM is primarily a representation-learning and visualisation method, not simply "an MLP trained without labels." Reading it as one of three related-but-distinct jobs keeps the comparison honest:
| Job | What it asks |
|---|---|
| Clustering | Which examples belong together? |
| Dimensionality reduction | Can this be described with far fewer numbers, without losing what matters? |
| Visualisation | Can this be laid out so a person can see its structure directly? |
A SOM does a version of all three at once — its units act like cluster centres, its 2D grid is a deliberate dimensionality reduction from whatever dimension the input had, and the grid itself is viewable directly — but it has no supervised classification objective built in anywhere; any class boundary drawn on top of a trained map is a separate step added afterward, not something the SOM's own training touches.
Related Topics
- Unsupervised Learning: Clustering and Dimensionality Reduction — k-means and PCA, covered in full, as the two more common methods for the same two underlying jobs a SOM blends together.
- Radial Basis Function Networks — another architecture built from a set of prototype/centre vectors, trained very differently (competitively here, by direct least-squares fit there).
- Concept Mapping and SOMs — this site's own use of SOMs specifically for laying out concept-space, a worked application of the method covered here.
References
Kohonen, T. (1982). Self-organized formation of topologically correct feature maps. Biological Cybernetics, 43(1), 59–69. ↩