Last updated: 2026-10-07

U
Undergraduate level
FDN
Foundational — Knowledge that endures for decades — core principles

Weightless Neural Networks and RAM Networks

Every model so far in this section stores what it's learned as numerical weights. RAM networks — also called weightless neural networks, or N-tuple recognisers — store it a completely different way: as values written directly into addressed memory, the same "RAM" as in a computer's random-access memory, not a claim about accessing inputs randomly. Where a weighted neuron multiplies, sums, and squashes, a RAM neuron does none of that arithmetic at all.

Weighted neuronRAM neuron
OperationMultiply inputs by weights, sum, activateForm an address from the inputs, look up what's stored there
What training changesWeight valuesContents of specific memory locations
Core arithmeticMultiplication and additionTable lookup and counting

An Elementary RAM Node FoundationalKnowledge that endures for decades — core principles

Take four binary inputs, say 1 0 1 1. Read as a binary number, that's an address: 1011₂ = 11₁₀. An elementary RAM node is nothing more than a small memory with one cell per possible address — 16 cells, for 4 binary inputs, since there are \(2^4\) possible 4-bit patterns. Training writes a value at the address the current input forms; recognition later looks up whatever is stored at that same address. The simplest version writes a plain 1 (seen before) versus 0 (never seen) — but richer variants exist and matter in practice: an integer count of how many times that pattern was seen during training, a confidence score, class-specific evidence when multiple classes share the same memory, or a thresholded ("bleached") count that only counts as a hit once the stored value clears some minimum.

Tuples and Discriminators FoundationalKnowledge that endures for decades — core principles

A real input is rarely just 4 bits. A larger binary pattern is split into fixed-size chunks — tuples — each one handled by its own RAM unit:

Complete pattern:  1 0 1 1 0 1 0 0
Tuple 1: 1011 -> RAM unit A
Tuple 2: 0100 -> RAM unit B

A class discriminator is a whole bank of RAM units, one per tuple, covering the entire input pattern for one particular class; recognising a new input means querying every unit in every class's discriminator and comparing how many of them respond. This is the WiSARD architecture, after Aleksander, Thomas, and Bowden's 1984 design, and it's also why RAM networks are sometimes called N-tuple recognisers — each RAM unit handles one N-bit tuple of the larger pattern1.WiSARD: WIlkie, Stonham and Aleksander's Recognition Device

Why It Generalises at All FoundationalKnowledge that endures for decades — core principles

A single RAM unit does not generalise — it only recognises the exact tuple patterns it was shown. Treating the whole network as equally rigid is the mistake to avoid: a complete input the network has never seen before can still be recognised correctly, because its individual tuples may each have appeared before, in other training examples, even though the full pattern never did. A new handwritten digit "3," for instance, may share several local tuple patterns with other training examples of "3" (a particular curve fragment, a particular stroke junction) without matching any single training example in full — recognition is built from a vote across many small, locally-learned fragments, not from matching one whole remembered template.

How well that generalisation actually works is governed by a short list of concrete design choices, not by chance:

  • Tuple size — larger tuples are more discriminating (fewer unrelated patterns collide at the same address) but need exponentially more memory, \(2^n\) addresses for an \(n\)-bit tuple, and reuse past observations less often.
  • Input mapping — which bits of the original pattern get grouped into which tuple, usually randomised once and then fixed, so that a local change to the input doesn't always land in the same single tuple.
  • Number of RAM units and aggregation rule — how many tuples cover the pattern, and how their individual responses are combined into one class score.
  • Training density, thresholds, and bleaching — how much training data was seen, and whether a hit has to clear a minimum stored count before it counts as recognition at all, which trades false positives against missed recognitions.
Note well. A RAM network's generalisation comes from tuples being individually reused across different whole inputs, not from "fuzzy" or approximate memory. Each tuple match is exact; what generalises is which combinations of exact local matches add up to a class decision.

Why This Is Attractive in Hardware Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

A simple RAM network can train in a single presentation of each example — write the pattern, move on — with no iterative gradient search and no multiplication anywhere in training or inference. That makes the approach attractive for low-power and embedded hardware, where table lookups and counters are cheap and floating-point multiplication is comparatively expensive, even though the encoding, tuple-mapping, and memory-sizing choices above still carry real design and computational cost of their own.

  • What Is a Learning Model? — this page's "what does training change" row (addressed memory values) contrasted directly with every weighted model's row (weights).
  • Nearest-Neighbour and Instance-Based Learning — another family that stores something other than weights, worth contrasting tuple-based local matching against whole-example distance matching.

References


  1. Aleksander, I., Thomas, W. V., & Bowden, P. A. (1984). WISARD: a radical step forward in image recognition. Sensor Review, 4(3), 120–124. ↩