Last updated: 2026-10-07

U
Undergraduate level
FDN
Foundational — Knowledge that endures for decades — core principles

How Weighted Neural Networks Learn

A hidden layer fixes what a single perceptron can represent, but it opens a new problem: the only thing directly observable during training is the error at the very end, at the output layer, yet every weight in every layer needs to know how much it personally contributed to that error before it can be corrected. The chain that answers this runs the same way for every weighted network in this section, from a two-layer toy example to a modern transformer:

graph LR A["Prediction"] --> B["Loss"] B --> C["Gradient"] C --> D["Parameter update"] D --> A

Loss: Turning "Wrong" Into a Number FoundationalKnowledge that endures for decades — core principles

A loss function takes the network's prediction and the true label and returns a single number measuring how bad that prediction was — zero for a perfect prediction, larger the further off it is. Mean squared error, \((y_{\text{true}} - y_{\text{pred}})^2\), is the natural choice for regression; cross-entropy, which penalises a confident wrong answer far more harshly than a hesitant one, is standard for classification. Whichever is used, training reduces to one search: find the weights that make this number as small as possible, averaged over the training set.

Gradient Descent FoundationalKnowledge that endures for decades — core principles

The gradient of the loss with respect to a weight is the answer to "if I nudge this one weight up very slightly, does the loss go up or down, and by how much?" Gradient descent uses that answer directly: move every weight a small step in whichever direction reduces the loss, repeat.

\[ w_i \leftarrow w_i - \eta\,\frac{\partial L}{\partial w_i} \]

The learning rate \(\eta\) sets how large that step is. Too large, and the updates overshoot the bottom of the loss and oscillate or diverge; too small, and training crawls toward the minimum so slowly it may never practically arrive, or gets stuck in a shallow local dip in the loss surface that a bigger step would have escaped. A full pass through the entire training set is one epoch; in practice, the gradient is estimated from a small random batch of examples at a time rather than the whole set at once, which is both cheaper per step and, somewhat counter-intuitively, often finds better solutions — the noise from using a different sample each step helps the search avoid settling into some of the shallower local dips.too big overshoots, too small crawls

Backpropagation: Sharing the Blame FoundationalKnowledge that endures for decades — core principles

Computing the gradient for every weight in a multi-layer network by brute force would mean separately asking "what if I nudge this weight" for every single weight — workable for a handful of parameters, hopeless for the millions a real network has. Backpropagation, popularised by Rumelhart, Hinton, and Williams in 1986, computes all of them in one efficient backward pass instead1. It works backward from the output layer, using the chain rule from calculus to relate the error at the output to the error at each hidden layer behind it, one layer at a time — each layer's share of the blame worked out from the next layer's, rather than computed from scratch. Stripped of the calculus: a weight whose output barely affected the final error gets barely nudged; a weight whose output strongly affected it, in a direction that made things worse, gets nudged hard in the opposite direction.

Note well. Backpropagation is an efficient way to compute a gradient that was already implicitly defined by the network and the loss function — it's bookkeeping, not a separate learning principle. The actual learning principle is still just gradient descent.

Overfitting and Underfitting FoundationalKnowledge that endures for decades — core principles

A network with enough weights can, in principle, drive the training loss to nearly zero by memorising the training set's specific quirks rather than the pattern underneath it — overfitting, visible as training loss that keeps falling while loss on held-out validation data stops improving or gets worse. Underfitting is the opposite failure: a network too small, or stopped too early, to capture the pattern at all, with both training and validation loss staying stubbornly high. Watching both curves during training, not just one, is how the two are told apart in practice.

References


  1. Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536. ↩