Last updated: 2026-10-07
Convolutional Neural Networks
A fully-connected layer treats every pixel of an image as an independent input, with its own separate weight to every hidden unit — which throws away the one fact that actually matters about images: a pattern worth detecting (an edge, a corner, a patch of colour) looks the same wherever in the image it appears. A Convolutional Neural Network (CNN) is built around that fact directly, through three ideas that work together: local connectivity, shared weights, and position-independent detection.
A Filter, by Hand FoundationalKnowledge that endures for decades — core principles
Take a small 4×4 grid of pixel values and a 2×2 filter (or kernel) designed to respond to a vertical edge:
Image (4x4): Filter (2x2):
0 0 1 1 1 -1
0 0 1 1 1 -1
0 0 1 1
0 0 1 1
Sliding the filter across the image and, at each position, multiplying it element-by-element against whatever 2×2 patch it currently covers and summing the result, is convolution. At the position covering the grid's top-left 2×2 patch (all zeros), the result is \(0\times1 + 0\times{-1} + 0\times1 + 0\times{-1} = 0\) — no edge there. At the position straddling the boundary between the 0s and the 1s, the result is strongly non-zero — the filter has found exactly the edge it was built to detect. Repeating this at every position produces a feature map: a smaller grid recording how strongly that one filter responded at every location in the original image.
Shared Weights, Not a Weight Per Pixel FoundationalKnowledge that endures for decades — core principles
The critical point is that the same four numbers — the filter — are reused at every single position, rather than each position getting its own independently-learned weights. The network learns one edge detector, not thousands of position-specific ones, and whatever it learns about detecting an edge in the top-left corner during training is automatically available for detecting the same edge anywhere else in a new image, without ever having seen that exact edge in that exact location before. A real CNN layer learns many filters in parallel (edges in several orientations, colour contrasts, simple textures), each producing its own feature map, stacked together as that layer's output.
Pooling and Hierarchical Features FoundationalKnowledge that endures for decades — core principles
Pooling — commonly taking the maximum value within each small region of a feature map — periodically shrinks the spatial resolution, giving the network some tolerance to a detected feature shifting by a few pixels between one image and the next, and reducing how much computation later layers need. Stacking several convolution-and-pooling layers builds a hierarchy: early layers' filters respond to simple edges and colour blobs; because later layers operate on the feature maps those early layers produced, not on raw pixels, they can combine simple features into textures and parts, and the layers after that combine parts into whole objects — a structure the network discovers through training, not one designed in by hand.
Related Topics
- Feedforward Networks and Multilayer Perceptrons — the fully-connected architecture this page's local, shared filters are a deliberate alternative to.
- Neural Network Architectures: From Perceptrons to Transformers — the full CNN architecture (LeCun et al., cited there in full), stride, padding, and channels, covered at greater depth.
- Computer Vision and Object Recognition — what these learned features replaced: hand-designed feature extraction.