Last updated: 2026-10-07
Radial Basis Function Networks
Every hidden unit in a Radial Basis Function (RBF) network asks the same simple question about a new input: how close is it to one particular remembered point? That remembered point is the unit's centre, and "how close" is measured by a response that peaks exactly at the centre and fades smoothly the further away the input gets — most commonly a Gaussian:
\[ \phi(\mathbf{x}) = \exp\!\left(-\frac{\lVert \mathbf{x} - \mathbf{c} \rVert^2}{2\sigma^2}\right) \]where \(\mathbf{c}\) is the centre and \(\sigma\) (the spread or width) controls how quickly the response falls off with distance. Broomhead and Lowe's 1988 paper established the network built from these units as a practical tool for function approximation1.
Local Versus Distributed Representation FoundationalKnowledge that endures for decades — core principles
This is the single biggest conceptual difference between an RBF network and the multilayer perceptron covered on the previous page. An MLP's hidden units each respond across a broad region of input space — their activation only goes to zero far from a boundary, not far from any particular point — which is a distributed representation: many units contribute a little to almost every prediction. An RBF unit responds strongly only near its own centre and is essentially silent everywhere else, a local representation: for any given input, only the handful of units whose centres happen to be nearby contribute anything at all. Local representations tend to train faster and behave more predictably near the centres actually covered by training data, at the cost of poor extrapolation anywhere no centre was placed nearby.one hidden unit, one neighbourhood
Centres, Output Weights, and Interpolation FoundationalKnowledge that endures for decades — core principles
A full RBF network has exactly one hidden layer of these units, followed by a simple weighted sum (no further nonlinearity needed) that combines their responses into the final output:
\[ y(\mathbf{x}) = \sum_i w_i\, \phi_i(\mathbf{x}) \]Fitting one, in the classic formulation, is two separate steps rather than one combined gradient search: first choose the centres (often simply a sample of the training points themselves, or the result of clustering the training data, covered on Unsupervised Learning), then solve directly for the output weights \(w_i\) that best fit the training labels given those fixed centres — a comparatively simple linear least-squares problem, not an iterative search, once the centres are fixed. Because each centre anchors a local bump of influence, the fitted function naturally interpolates smoothly between the centres it has: near any training point, the matching centre dominates and the output is close to that point's label; between two training points, the two nearest centres' bumps blend.
Related Topics
- Feedforward Networks and Multilayer Perceptrons — the distributed-representation contrast this page's local units are defined against.
- Nearest-Neighbour and Instance-Based Learning — a different, even more local method, which skips fitting a smooth response entirely and just votes among whichever stored examples are nearest.
- Unsupervised Learning: Clustering and Dimensionality Reduction — a common source for this page's centres, when they aren't simply the training points themselves.
References
Broomhead, D. S., & Lowe, D. (1988). Multivariable functional interpolation and adaptive networks. Complex Systems, 2(3), 321–355. ↩