Last updated: 2026-10-07
Feedforward Networks and Multilayer Perceptrons
Stack artificial neurons into layers — an input layer, one or more hidden layers, an output layer — with every unit in one layer feeding every unit in the next, and no connection ever running backward during inference, and the result is a multilayer perceptron (MLP), the direct architectural fix for the single perceptron's linear-only limitation. Information moves strictly forward, input to output, which is what "feedforward" names.
Depth, Width, and Nonlinearity FoundationalKnowledge that endures for decades — core principles
Width is how many units sit in a given layer; depth is how many layers the input passes through on its way to the output. Both add representational capacity, but only if the hidden units' activation functions are themselves nonlinear — stacking purely linear layers collapses algebraically back into one single linear layer, since a linear function of a linear function is still just linear, gaining nothing from the extra depth. This is why every hidden layer in a genuinely useful MLP applies a nonlinear activation (sigmoid, ReLU, or similar) after its weighted sum, not just the output layer.
Universal Approximation, Read Correctly FoundationalKnowledge that endures for decades — core principles
The universal approximation theorem, proven by Hornik, Stinchcombe, and White in 1989, states that a feedforward network with even a single hidden layer, given enough units in it, can approximate any continuous function on a bounded input region to arbitrary precision1. It's a genuine, proven result, and it is routinely overstated. It guarantees that a suitable network exists — it says nothing about how many units that might take (sometimes astronomically many, for a single layer, versus a modest number spread across several), nothing about whether gradient descent can actually find those weights from a random starting point, and nothing about how much training data would be needed to pin them down. In practice, networks with several narrower layers routinely outperform one enormously wide layer at the same total parameter count, which is a large part of why "deep" learning means more layers, not just more units in one.exists, not: is efficiently learnable
Classification Versus Regression Outputs FoundationalKnowledge that endures for decades — core principles
The same hidden layers can feed either kind of task; only the output stage differs. A regression output is typically a single unit with no activation at all (or a linear one), passing the raw weighted sum straight through as the predicted number. A classification output over \(k\) classes typically uses \(k\) units and a softmax activation, which turns the raw scores into a proper probability distribution — all values between 0 and 1, summing to exactly 1 — so "70% confident this is class A" has an actual probabilistic meaning rather than being an arbitrary scaled score.
Related Topics
- The Perceptron and Linear Classification — the single-layer case, and the XOR limitation this architecture exists to fix.
- How Weighted Neural Networks Learn — the training method (loss, gradient descent, backpropagation) that fits this architecture's weights.
- Neural Network Architectures: From Perceptrons to Transformers — where this page's feedforward building block gets extended with convolution, recurrence, and attention, covered in full.
References
Hornik, K., Stinchcombe, M., & White, H. (1989). Multilayer feedforward networks are universal approximators. Neural Networks, 2(5), 359–366. ↩