Last updated: 2026-10-07

U
Undergraduate level
FDN
Foundational — Knowledge that endures for decades — core principles

Recurrent Neural Networks

Every architecture so far treats each input as a single, self-contained snapshot. A sentence, an audio clip, a stock price history — these have sequential structure, where the right answer at one position usually depends on what came before it, and a fixed-size feedforward network has no natural place to keep that context. A recurrent neural network (RNN) answers this by feeding its own previous output back in as part of its current input:

\[ h_t = f(W_x x_t + W_h h_{t-1} + b) \]

\(h_t\), the hidden state at sequence position \(t\), depends on both the current input \(x_t\) and the previous hidden state \(h_{t-1}\) — which itself depended on the one before it, and so on back to the start of the sequence. The same weights \(W_x\) and \(W_h\) are reused at every position; "unfolding" the recurrence across a short sequence makes this visible as a chain of identical units, each handing its hidden state to the next.

graph LR X1["x1"] --> H1["h1"] H1 --> H2["h2"] X2["x2"] --> H2 H2 --> H3["h3"] X3["x3"] --> H3 H3 --> O["output"]

What the Hidden State Actually Is FoundationalKnowledge that endures for decades — core principles

Calling \(h_t\) "memory" is common shorthand, and slightly misleading if taken literally. It's a learned state representation shaped entirely by training to be useful for the task at hand — it can carry information forward across many steps, but it can just as easily forget it, distort it, or overwrite it with something more locally useful, because nothing in the basic equation above explicitly protects any piece of information from being squashed by the next update. Whether a given RNN actually preserves some particular signal long enough to matter is an empirical question about that trained network, not a guarantee the architecture provides.it can forget, distort, or overwrite, not just store

Many-to-One and Many-to-Many FoundationalKnowledge that endures for decades — core principles

The same recurrence supports several different task shapes, depending on when the output is read out: a many-to-one task (classify a whole sentence's sentiment) reads only the final hidden state, after the whole sequence has been processed; a many-to-many task (label every word in a sentence, or translate one sequence into another) reads an output at every step, or produces a second sequence once the first has been fully consumed.

The Vanishing Gradient, and Gated Fixes Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

Training an RNN runs backpropagation back through every time step the way an ordinary network runs it back through every layer — and a sequence of 1,000 steps is, for this purpose, a network 1,000 layers deep. Repeatedly multiplying a gradient signal by the same recurrent weights as it's pushed back through that many steps tends to shrink it toward zero (or, less commonly, blow it up) well before it reaches the early steps that may have mattered most — the vanishing gradient problem, and the reason plain RNNs are notoriously bad at long-range dependencies in practice, however clean the equation looks on paper. Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) architectures address this with learned gates — small sub-networks that decide what to write into the hidden state, what to read out of it, and what to actively forget — giving the network an explicit, trainable mechanism for preserving a signal across many steps, rather than hoping it happens to survive repeated multiplication. Gating mitigates the vanishing-gradient problem; it does not make it disappear under all conditions, and an LSTM trained on a genuinely very long sequence can still lose information it would have been useful to keep.