Last updated: 2026-10-07
What Is a Learning Model?
Every page in this section describes a different way of answering the same question: given some examples, how should a system change itself so that it handles the next one correctly? The honest answer is never "it thinks about it" — it's always some specific, mechanical change to some specific, persistent piece of internal state. A perceptron's answer is to adjust a set of weights. A decision tree's answer is to grow a new branch. A nearest-neighbour classifier's answer is to simply remember the example. Those are three genuinely different mechanisms, not three names for the same thing, and keeping that difference in view is the one habit that makes the rest of this section make sense.
The Vocabulary FoundationalKnowledge that endures for decades — core principles
A small, concrete example carries the rest of this page, and recurs on every page that follows it. Ten points, each described by two measurements, each belonging to one of two classes:
| Feature 1 | Feature 2 | Class |
|---|---|---|
| 1.0 | 1.5 | round |
| 1.5 | 1.0 | round |
| 2.0 | 2.0 | round |
| 1.0 | 2.5 | round |
| 2.5 | 1.5 | round |
| 4.0 | 4.0 | square |
| 4.5 | 3.5 | square |
| 3.5 | 4.5 | square |
| 5.0 | 4.5 | square |
| 4.0 | 5.0 | square |
The vocabulary below is standard across the field and follows the treatment in Hastie, Tibshirani, and Friedman's standard reference1, already cited in full on this site's own supervised-learning page. Each row of the table is an example (sometimes called an observation or an instance). The two numbers are its features — measured properties that might help tell the classes apart, chosen in advance by whoever built the dataset, not discovered by the model. "Round" and "square" are labels, in this case a classification task, because the thing being predicted is one of a fixed set of categories rather than a number on a continuous scale; predicting a number — a price, a temperature, a count — is regression instead, and most of what follows applies to both with the output stage swapped.
A model is whatever internal structure gets built or adjusted from the examples — a set of weights, a tree, a stored list of points, a table of memory cells. The process of building or adjusting it is training, and using the finished model to produce an answer for a new example it hasn't seen before is inference. Training almost always has a target to aim at: a loss (or error) that measures how wrong the model's current predictions are, computed by comparing its output against the known labels. Most of training, across every family in this section, is some version of "change the model's state in whatever direction reduces the loss" — what changes, and how, is exactly the distinction this section is built around.a model is the thing that changed, not the algorithm that changed it
This is supervised learning: every training example already carries its correct label, and the model is learning a mapping from inputs to those known outputs. Unsupervised learning gets no labels at all, and asks a different question — what structure is already present in the data, with no "correct answer" to check against — covered in full on Unsupervised Learning: Clustering and Dimensionality Reduction. Self-supervised learning sits between the two: it manufactures its own labels from the structure of the data itself (predict the next word, reconstruct a corrupted input) rather than needing a human to supply them. Everything in this section is supervised unless a page says otherwise.
Why a Held-Out Test Set FoundationalKnowledge that endures for decades — core principles
A model that has only ever been checked against the data it trained on can look perfect and still be useless, because nothing stops it from simply memorising that specific data rather than learning the pattern that generated it. The standard discipline against this is to split the available examples into three groups before training starts: a training set the model actually learns from, a validation set used to compare different models or settings against each other during development, and a test set, touched exactly once, at the very end, to report how the final choice actually performs. Generalisation is the name for a model's performance on genuinely new examples — the test set is the only honest way to measure it, precisely because it was never available to influence any decision made while building the model.
What Gets Changed FoundationalKnowledge that endures for decades — core principles
This is the theme the rest of the section develops one family at a time. "Training" sounds like one activity, but what training actually writes differs sharply between model families, and that difference predicts a great deal about each one's behaviour — how much data it needs, how it fails, whether it can be updated one example at a time, what it means for it to "remember" something.
| Model family | What training changes |
|---|---|
| Perceptron / weighted neural network | A set of numerical weights and biases |
| RAM / weightless network | Values written directly into addressed memory cells |
| Decision tree | A growing set of branch tests |
| Support vector machine | Which training points count as the boundary's supports |
| Nearest-neighbour model | Nothing — it keeps the examples themselves |
None of these is "more correct" than the others in the abstract — each is a different bet about what kind of structure is worth keeping, and each pays off on a different kind of problem. The rest of this section works through what each bet actually buys.
Related Topics
- What Is an Artificial Neuron? — the first model family this section develops, and the vocabulary above applied to a concrete mechanism.
- Supervised Learning: Regression, Trees, and Ensembles — decision trees, ensembles, and the precision/recall vocabulary for judging a classifier, covered in full.
- Unsupervised Learning: Clustering and Dimensionality Reduction — the label-free counterpart to everything on this page.
- Symbolic and Connectionist Reasoning: Two Ways to Build "Think" — the broader distinction this section's own family-by-family contrasts sit inside.
References
Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction (2nd ed.). Springer. Held by the University of Reading Library. ↩