Last updated: 2026-10-07
How Weighted Neural Networks Learn
A hidden layer fixes what a single perceptron can represent, but it opens a new problem: the only thing directly observable during training is the error at the very end, at the output layer, yet every weight in every layer needs to know how much it personally contributed to that error before it can be corrected. The chain that answers this runs the same way for every weighted network in this section, from a two-layer toy example to a modern transformer:
Loss: Turning "Wrong" Into a Number FoundationalKnowledge that endures for decades — core principles
A loss function takes the network's prediction and the true label and returns a single number measuring how bad that prediction was — zero for a perfect prediction, larger the further off it is. Mean squared error, \((y_{\text{true}} - y_{\text{pred}})^2\), is the natural choice for regression; cross-entropy, which penalises a confident wrong answer far more harshly than a hesitant one, is standard for classification. Whichever is used, training reduces to one search: find the weights that make this number as small as possible, averaged over the training set.
Gradient Descent FoundationalKnowledge that endures for decades — core principles
The gradient of the loss with respect to a weight is the answer to "if I nudge this one weight up very slightly, does the loss go up or down, and by how much?" Gradient descent uses that answer directly: move every weight a small step in whichever direction reduces the loss, repeat.
\[ w_i \leftarrow w_i - \eta\,\frac{\partial L}{\partial w_i} \]The learning rate \(\eta\) sets how large that step is. Too large, and the updates overshoot the bottom of the loss and oscillate or diverge; too small, and training crawls toward the minimum so slowly it may never practically arrive, or gets stuck in a shallow local dip in the loss surface that a bigger step would have escaped. A full pass through the entire training set is one epoch; in practice, the gradient is estimated from a small random batch of examples at a time rather than the whole set at once, which is both cheaper per step and, somewhat counter-intuitively, often finds better solutions — the noise from using a different sample each step helps the search avoid settling into some of the shallower local dips.too big overshoots, too small crawls
Backpropagation: Sharing the Blame FoundationalKnowledge that endures for decades — core principles
Computing the gradient for every weight in a multi-layer network by brute force would mean separately asking "what if I nudge this weight" for every single weight — workable for a handful of parameters, hopeless for the millions a real network has. Backpropagation, popularised by Rumelhart, Hinton, and Williams in 1986, computes all of them in one efficient backward pass instead1. It works backward from the output layer, using the chain rule from calculus to relate the error at the output to the error at each hidden layer behind it, one layer at a time — each layer's share of the blame worked out from the next layer's, rather than computed from scratch. Stripped of the calculus: a weight whose output barely affected the final error gets barely nudged; a weight whose output strongly affected it, in a direction that made things worse, gets nudged hard in the opposite direction.
Overfitting and Underfitting FoundationalKnowledge that endures for decades — core principles
A network with enough weights can, in principle, drive the training loss to nearly zero by memorising the training set's specific quirks rather than the pattern underneath it — overfitting, visible as training loss that keeps falling while loss on held-out validation data stops improving or gets worse. Underfitting is the opposite failure: a network too small, or stopped too early, to capture the pattern at all, with both training and validation loss staying stubbornly high. Watching both curves during training, not just one, is how the two are told apart in practice.
Related Topics
- The Perceptron and Linear Classification — the single-layer case this page's hidden-layer credit-assignment problem builds on.
- Feedforward Networks and Multilayer Perceptrons — the architecture this training method is actually run against.
- Calculus and Optimization — the full derivative mechanics behind the gradient this page treats conceptually.
References
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536. ↩