Last updated: 2026-10-07
The Perceptron and Linear Classification
Rosenblatt's perceptron, from 1958, is the first complete learning algorithm this section covers: a single artificial neuron with a threshold activation, plus a rule for adjusting its weights automatically from labelled examples, rather than setting them by hand1. What it draws is always the same kind of object — a straight line in two dimensions, a flat plane in three, a hyperplane in general — defined by:
\[ \mathbf{w}^\mathsf{T}\mathbf{x} + b = 0 \]Every point on one side gets classified one way, every point on the other side the other way, and the weights and bias set exactly where that dividing line sits and which way it tilts.
The Learning Rule FoundationalKnowledge that endures for decades — core principles
Training a perceptron means repeatedly cycling through the labelled examples and, for each one, checking whether the current weights get it right. If they do, nothing changes. If they don't, the weights shift a little in the direction that would have made that example's prediction correct:
\[ w_i \leftarrow w_i + \eta\,(y_{\text{true}} - y_{\text{pred}})\,x_i \]where \(\eta\) is a small learning rate. Run against this section's shared dataset — five "round" points clustered near the origin, five "square" points clustered further out — the rule converges in only a handful of passes, because a single straight line genuinely separates the two clusters cleanly. The perceptron convergence theorem guarantees this in general: if the data is linearly separable at all, this update rule is guaranteed to find a separating line in a finite number of steps. It says nothing about data that isn't separable, which is exactly where the trouble starts.
The XOR Problem FoundationalKnowledge that endures for decades — core principles
XOR — output 1 when exactly one of two binary inputs is 1, output 0 otherwise — is the simplest function a single perceptron provably cannot learn. Plot its four input/output pairs on a 2D grid and the two true outputs sit on opposite corners from each other: no single straight line has both true points on one side and both false points on the other, no matter how the weights are set.
| x1 | x2 | AND | OR | XOR |
|---|---|---|---|---|
| 0 | 0 | 0 | 0 | 0 |
| 1 | 0 | 0 | 1 | 1 |
| 0 | 1 | 0 | 1 | 1 |
| 1 | 1 | 1 | 1 | 0 |
AND and OR are both linearly separable — a single line handles each, the way the artificial-neuron page's worked example handled AND directly. XOR is not, and Minsky and Papert's 1969 analysis of exactly this limitation is widely credited with cooling funding for neural-network research through the 1970s2.a published proof, not just an observation
The Way Out: a Hidden Layer FoundationalKnowledge that endures for decades — core principles
XOR becomes separable the moment the input is re-expressed in different coordinates first. Feed \((x_1, x_2)\) through two hidden units — one computing roughly "\(x_1\) OR \(x_2\)", the other roughly "\(x_1\) AND \(x_2\)" — and XOR is just "first is true, second is false," which a third, output-layer perceptron separates with a single straight line in that two-dimensional space, even though no such line existed in the original input space. Nothing about the individual units changed; stacking them changed what coordinates the final decision gets to be drawn in. That's the idea the next page develops into a full training method, because a hidden layer only helps if there's a way to train its weights too, and the credit for an error now has to be shared across two layers rather than read directly off one.
Related Topics
- What Is an Artificial Neuron? — the single-unit mechanism this page trains for the first time.
- Neural Network Architectures: From Perceptrons to Transformers — covers the same XOR limitation and names the same Minsky and Papert result, from the "lineage of fixes" framing this section deliberately sets alongside rather than replaces.