Last updated: 2026-10-07

U
Undergraduate level
FDN
Foundational — Knowledge that endures for decades — core principles

Feedforward Networks and Multilayer Perceptrons

Stack artificial neurons into layers — an input layer, one or more hidden layers, an output layer — with every unit in one layer feeding every unit in the next, and no connection ever running backward during inference, and the result is a multilayer perceptron (MLP), the direct architectural fix for the single perceptron's linear-only limitation. Information moves strictly forward, input to output, which is what "feedforward" names.

graph LR I1["input"] --> H1["hidden"] I2["input"] --> H1 I1 --> H2["hidden"] I2 --> H2 H1 --> O["output"] H2 --> O

Depth, Width, and Nonlinearity FoundationalKnowledge that endures for decades — core principles

Width is how many units sit in a given layer; depth is how many layers the input passes through on its way to the output. Both add representational capacity, but only if the hidden units' activation functions are themselves nonlinear — stacking purely linear layers collapses algebraically back into one single linear layer, since a linear function of a linear function is still just linear, gaining nothing from the extra depth. This is why every hidden layer in a genuinely useful MLP applies a nonlinear activation (sigmoid, ReLU, or similar) after its weighted sum, not just the output layer.

Universal Approximation, Read Correctly FoundationalKnowledge that endures for decades — core principles

The universal approximation theorem, proven by Hornik, Stinchcombe, and White in 1989, states that a feedforward network with even a single hidden layer, given enough units in it, can approximate any continuous function on a bounded input region to arbitrary precision1. It's a genuine, proven result, and it is routinely overstated. It guarantees that a suitable network exists — it says nothing about how many units that might take (sometimes astronomically many, for a single layer, versus a modest number spread across several), nothing about whether gradient descent can actually find those weights from a random starting point, and nothing about how much training data would be needed to pin them down. In practice, networks with several narrower layers routinely outperform one enormously wide layer at the same total parameter count, which is a large part of why "deep" learning means more layers, not just more units in one.exists, not: is efficiently learnable

Classification Versus Regression Outputs FoundationalKnowledge that endures for decades — core principles

The same hidden layers can feed either kind of task; only the output stage differs. A regression output is typically a single unit with no activation at all (or a linear one), passing the raw weighted sum straight through as the predicted number. A classification output over \(k\) classes typically uses \(k\) units and a softmax activation, which turns the raw scores into a proper probability distribution — all values between 0 and 1, summing to exactly 1 — so "70% confident this is class A" has an actual probabilistic meaning rather than being an arbitrary scaled score.

References


  1. Hornik, K., Stinchcombe, M., & White, H. (1989). Multilayer feedforward networks are universal approximators. Neural Networks, 2(5), 359–366. ↩