Last updated: 2026-10-07

U
Undergraduate level
FDN
Foundational — Knowledge that endures for decades — core principles

What Is a Learning Model?

Every page in this section describes a different way of answering the same question: given some examples, how should a system change itself so that it handles the next one correctly? The honest answer is never "it thinks about it" — it's always some specific, mechanical change to some specific, persistent piece of internal state. A perceptron's answer is to adjust a set of weights. A decision tree's answer is to grow a new branch. A nearest-neighbour classifier's answer is to simply remember the example. Those are three genuinely different mechanisms, not three names for the same thing, and keeping that difference in view is the one habit that makes the rest of this section make sense.

The Vocabulary FoundationalKnowledge that endures for decades — core principles

A small, concrete example carries the rest of this page, and recurs on every page that follows it. Ten points, each described by two measurements, each belonging to one of two classes:

Feature 1Feature 2Class
1.01.5round
1.51.0round
2.02.0round
1.02.5round
2.51.5round
4.04.0square
4.53.5square
3.54.5square
5.04.5square
4.05.0square

The vocabulary below is standard across the field and follows the treatment in Hastie, Tibshirani, and Friedman's standard reference1, already cited in full on this site's own supervised-learning page. Each row of the table is an example (sometimes called an observation or an instance). The two numbers are its features — measured properties that might help tell the classes apart, chosen in advance by whoever built the dataset, not discovered by the model. "Round" and "square" are labels, in this case a classification task, because the thing being predicted is one of a fixed set of categories rather than a number on a continuous scale; predicting a number — a price, a temperature, a count — is regression instead, and most of what follows applies to both with the output stage swapped.

A model is whatever internal structure gets built or adjusted from the examples — a set of weights, a tree, a stored list of points, a table of memory cells. The process of building or adjusting it is training, and using the finished model to produce an answer for a new example it hasn't seen before is inference. Training almost always has a target to aim at: a loss (or error) that measures how wrong the model's current predictions are, computed by comparing its output against the known labels. Most of training, across every family in this section, is some version of "change the model's state in whatever direction reduces the loss" — what changes, and how, is exactly the distinction this section is built around.a model is the thing that changed, not the algorithm that changed it

graph TD A["Training examples"] --> B["Learning process"] B --> C["Learned model"] D["New example"] --> C C --> E["Prediction"]

This is supervised learning: every training example already carries its correct label, and the model is learning a mapping from inputs to those known outputs. Unsupervised learning gets no labels at all, and asks a different question — what structure is already present in the data, with no "correct answer" to check against — covered in full on Unsupervised Learning: Clustering and Dimensionality Reduction. Self-supervised learning sits between the two: it manufactures its own labels from the structure of the data itself (predict the next word, reconstruct a corrupted input) rather than needing a human to supply them. Everything in this section is supervised unless a page says otherwise.

Why a Held-Out Test Set FoundationalKnowledge that endures for decades — core principles

A model that has only ever been checked against the data it trained on can look perfect and still be useless, because nothing stops it from simply memorising that specific data rather than learning the pattern that generated it. The standard discipline against this is to split the available examples into three groups before training starts: a training set the model actually learns from, a validation set used to compare different models or settings against each other during development, and a test set, touched exactly once, at the very end, to report how the final choice actually performs. Generalisation is the name for a model's performance on genuinely new examples — the test set is the only honest way to measure it, precisely because it was never available to influence any decision made while building the model.

Note well. A model's performance on its own training data is not a measurement of how good it is — it's a measurement of how much it has memorised. Generalisation can only be checked against examples the model never saw during training.

What Gets Changed FoundationalKnowledge that endures for decades — core principles

This is the theme the rest of the section develops one family at a time. "Training" sounds like one activity, but what training actually writes differs sharply between model families, and that difference predicts a great deal about each one's behaviour — how much data it needs, how it fails, whether it can be updated one example at a time, what it means for it to "remember" something.

Model familyWhat training changes
Perceptron / weighted neural networkA set of numerical weights and biases
RAM / weightless networkValues written directly into addressed memory cells
Decision treeA growing set of branch tests
Support vector machineWhich training points count as the boundary's supports
Nearest-neighbour modelNothing — it keeps the examples themselves

None of these is "more correct" than the others in the abstract — each is a different bet about what kind of structure is worth keeping, and each pays off on a different kind of problem. The rest of this section works through what each bet actually buys.

References


  1. Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction (2nd ed.). Springer. Held by the University of Reading Library. ↩