Last updated: 2026-10-07

U
Undergraduate level
FDN
Foundational — Knowledge that endures for decades — core principles

Neural Networks and Alternative Models

Laid out as a single historical line — perceptron, then MLP, then the rest — this section's architectures would look like one lineage of fixes. They aren't. Weighted networks, RAM networks, prototype-based networks, attractor networks, margin-based classifiers, and instance-based models are substantially different traditions that developed in parallel, each making a different bet about what's worth keeping from training data. This page collects them into several different views side by side, because any single taxonomy flattens distinctions a reader actually needs.

By What Gets Stored FoundationalKnowledge that endures for decades — core principles

This is the organising question the first page of this section opened with, and it sorts every architecture covered here into a small number of genuinely different categories.

Stored stateModels
Continuous weights, fitted by gradient descentPerceptron, MLP, RNN
Explicit, addressed memory contentsRAM / weightless networks
Prototype or centre vectorsRBF networks, Self-Organising Maps
Connection weights encoding attractor statesHopfield networks
A subset of the training examples themselves, chosen by marginSupport Vector Machines
Every training example, unmodifiedNearest-neighbour models

By How Information Moves FoundationalKnowledge that endures for decades — core principles

FlowModels
Feedforward — input to output, no cycles during inferencePerceptron, MLP, RBF network
Recurrent — output fed back as input, step by stepRNN
Recurrent until settled — feedback run to a stable point, not a fixed number of stepsHopfield network
Addressed lookup — no propagation at all, one table readRAM network
Neighbourhood competition — one winner, graded update to nearby unitsSelf-Organising Map

By How Training Actually Works FoundationalKnowledge that endures for decades — core principles

MechanismModels
Gradient-based optimisation of a loss functionPerceptron, MLP, RNN
Direct memory writes or count updatesRAM network
Competitive, neighbourhood-weighted updatesSelf-Organising Map
Centre selection, then a direct least-squares fitRBF network
Margin optimisation over a constrained set of candidatesSupport Vector Machine
No fitting at all — retain the data, decide at query timeNearest-neighbour model

No single one of these three views is more "correct" than the others — each answers a different question about the same set of models, and a real system occasionally combines ideas from more than one row at once (an RBF network, for instance, is feedforward in information flow but fits its output layer the way a linear regression does, not the way an MLP's hidden layers do). That's more truthful than a single evolutionary tree running from perceptrons to transformers, which is exactly the oversimplification this section was built to correct.three honest views, not one tidy family tree

A Working Comparison FoundationalKnowledge that endures for decades — core principles

ModelPersistent stateTypical learning processKey inductive bias
PerceptronWeights and biasError-driven updatesLinear separation
MLPLayer weights and biasesBackpropagationComposed nonlinear functions
RAM networkAddressed memory valuesDirect writes or count updatesRecognition from stored tuples
RBF networkCentres, widths, output weightsCentre selection, then least-squares fitLocal similarity
Self-Organising MapPrototype vectors on a gridCompetitive neighbourhood updatesTopological organisation
RNNWeights defining state transitionsBackpropagation through timeSequential dependency
Hopfield networkRecurrent connection weightsDirect pattern storageAttractor recall
Support Vector MachineSupport vectors and coefficientsMargin-maximising optimisationWide separating margin
k-Nearest-NeighbourStored examplesNone — defer to query timeNearby examples share an output

Real systems routinely combine several of these ideas rather than picking exactly one row — a modern image pipeline might use a convolutional network's learned features as the input to a much simpler nearest-neighbour or SVM classifier, for instance — so this table describes the pure, textbook version of each idea, not a claim that production systems always keep them this separate.