Last updated: 2026-10-07
Neural Networks and Alternative Models
Laid out as a single historical line — perceptron, then MLP, then the rest — this section's architectures would look like one lineage of fixes. They aren't. Weighted networks, RAM networks, prototype-based networks, attractor networks, margin-based classifiers, and instance-based models are substantially different traditions that developed in parallel, each making a different bet about what's worth keeping from training data. This page collects them into several different views side by side, because any single taxonomy flattens distinctions a reader actually needs.
By What Gets Stored FoundationalKnowledge that endures for decades — core principles
This is the organising question the first page of this section opened with, and it sorts every architecture covered here into a small number of genuinely different categories.
| Stored state | Models |
|---|---|
| Continuous weights, fitted by gradient descent | Perceptron, MLP, RNN |
| Explicit, addressed memory contents | RAM / weightless networks |
| Prototype or centre vectors | RBF networks, Self-Organising Maps |
| Connection weights encoding attractor states | Hopfield networks |
| A subset of the training examples themselves, chosen by margin | Support Vector Machines |
| Every training example, unmodified | Nearest-neighbour models |
By How Information Moves FoundationalKnowledge that endures for decades — core principles
| Flow | Models |
|---|---|
| Feedforward — input to output, no cycles during inference | Perceptron, MLP, RBF network |
| Recurrent — output fed back as input, step by step | RNN |
| Recurrent until settled — feedback run to a stable point, not a fixed number of steps | Hopfield network |
| Addressed lookup — no propagation at all, one table read | RAM network |
| Neighbourhood competition — one winner, graded update to nearby units | Self-Organising Map |
By How Training Actually Works FoundationalKnowledge that endures for decades — core principles
| Mechanism | Models |
|---|---|
| Gradient-based optimisation of a loss function | Perceptron, MLP, RNN |
| Direct memory writes or count updates | RAM network |
| Competitive, neighbourhood-weighted updates | Self-Organising Map |
| Centre selection, then a direct least-squares fit | RBF network |
| Margin optimisation over a constrained set of candidates | Support Vector Machine |
| No fitting at all — retain the data, decide at query time | Nearest-neighbour model |
No single one of these three views is more "correct" than the others — each answers a different question about the same set of models, and a real system occasionally combines ideas from more than one row at once (an RBF network, for instance, is feedforward in information flow but fits its output layer the way a linear regression does, not the way an MLP's hidden layers do). That's more truthful than a single evolutionary tree running from perceptrons to transformers, which is exactly the oversimplification this section was built to correct.three honest views, not one tidy family tree
A Working Comparison FoundationalKnowledge that endures for decades — core principles
| Model | Persistent state | Typical learning process | Key inductive bias |
|---|---|---|---|
| Perceptron | Weights and bias | Error-driven updates | Linear separation |
| MLP | Layer weights and biases | Backpropagation | Composed nonlinear functions |
| RAM network | Addressed memory values | Direct writes or count updates | Recognition from stored tuples |
| RBF network | Centres, widths, output weights | Centre selection, then least-squares fit | Local similarity |
| Self-Organising Map | Prototype vectors on a grid | Competitive neighbourhood updates | Topological organisation |
| RNN | Weights defining state transitions | Backpropagation through time | Sequential dependency |
| Hopfield network | Recurrent connection weights | Direct pattern storage | Attractor recall |
| Support Vector Machine | Support vectors and coefficients | Margin-maximising optimisation | Wide separating margin |
| k-Nearest-Neighbour | Stored examples | None — defer to query time | Nearby examples share an output |
Real systems routinely combine several of these ideas rather than picking exactly one row — a modern image pipeline might use a convolutional network's learned features as the input to a much simpler nearest-neighbour or SVM classifier, for instance — so this table describes the pure, textbook version of each idea, not a claim that production systems always keep them this separate.
Related Topics
- What Is a Learning Model? — where the "what gets stored" question that organises this whole page was first raised.
- Symbolic and Connectionist Reasoning: Two Ways to Build "Think" — a different, complementary cut across the same territory, one level up from architecture choice.