Last updated: 2026-10-07
Support Vector Machines
A Support Vector Machine (SVM) is a supervised, maximum-margin classifier — not, despite sharing this section with so many architectures that are, an artificial neural network at all. It belongs here for the same reason RAM networks and nearest-neighbour models do: it stores and learns in a genuinely different way than a weighted network does, and understanding that difference is the actual point of this section, not a historical footnote to it.
Many Lines Separate These Groups — Which One? FoundationalKnowledge that endures for decades — core principles
Given two linearly separable classes — this section's shared "round" and "square" dataset is a clean example — a perceptron stops the moment it finds any line that separates them, with no preference among the infinitely many lines that would all work equally well on the training data. Cortes and Vapnik's 1995 formulation asks a sharper question: among all the separating lines, which one leaves the widest possible corridor of empty space before it touches either class1? That corridor is the margin, and the SVM's decision boundary is the one line that maximises it.
Support Vectors: the Points That Matter FoundationalKnowledge that endures for decades — core principles
The training points that end up lying exactly on the edge of that margin — the closest point or points from each class — are the support vectors, and they are the only training points that actually determine where the boundary sits. Every other point could be moved further from the boundary, or removed from the training set entirely, without changing the fitted line at all. This is a genuinely distinctive property: a decision tree's split points and a neural network's weights are both shaped, in principle, by every single training example; an SVM's boundary is shaped by only the handful nearest to it.most training points are irrelevant to the final boundary
Soft Margins: When the Data Isn't Perfectly Separable FoundationalKnowledge that endures for decades — core principles
Real data is rarely as cleanly separable as the toy dataset above — a few outliers or mislabelled points can make a perfect separating line impossible, or force the margin so narrow it's useless. A soft margin allows some training points to sit inside the margin, or even on the wrong side of the boundary entirely, at a cost controlled by a regularisation parameter \(C\): a large \(C\) penalises every margin violation heavily, producing a narrower margin that fits the training data closely (and risks overfitting to its noise); a small \(C\) tolerates more violations in exchange for a wider, more conservative margin that generalises better when the training data is genuinely noisy.
Kernels: Nonlinear Boundaries Without Explicit Coordinates FoundationalKnowledge that endures for decades — core principles
A single straight margin only separates linearly separable classes, which is a real limitation — until the kernel trick is added. A kernel function computes something that behaves exactly like an inner product calculated in some other, usually much higher-dimensional feature space, without the SVM ever having to construct or store the actual coordinates of that space explicitly. A polynomial kernel behaves as though the original features had been expanded with every pairwise product and power up to some degree; a radial basis function kernel (the same Gaussian-shaped response covered on the RBF networks page) behaves as though every training point had been given its own bump of influence. Because the SVM's whole fitting procedure only ever needs those inner products, never the raw high-dimensional coordinates themselves, the trick is computationally practical even when the implied feature space would be far too large to construct directly — sometimes even infinite-dimensional, in the Gaussian kernel's case.
Support Vector Regression extends the same margin idea to predicting a continuous value rather than a class: instead of maximising the gap between two classes, it fits a tube of fixed width around the regression line and only penalises points that fall outside that tube, ignoring small errors within it entirely.
Related Topics
- The Perceptron and Linear Classification — the simpler linear classifier an SVM's margin-maximising boundary directly improves on.
- Radial Basis Function Networks — the same Gaussian response reused here as a kernel rather than as a hidden-layer unit.
- Nearest-Neighbour and Instance-Based Learning — another model whose prediction depends only on a subset of stored training points, worth contrasting against which subset and why.
References
Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297. ↩