Last updated: 2026-10-07
Attention and Transformers
This page assumes the RNN page's own starting question — what if the right output depends on earlier elements of a sequence — and picks up exactly where that page's own limitation left off: a plain recurrent network routes information through every intervening step to connect two distant positions, one step at a time, which is both slow to train and liable to lose a long-range signal along the way. A transformer, introduced by Vaswani and colleagues in 2017, removes the step-by-step chain entirely1.
Tokens and Embeddings FoundationalKnowledge that endures for decades — core principles
A sequence (a sentence, say) is first broken into tokens — words or word-fragments — and each token is converted into an embedding, a vector of numbers positioned so that tokens used in similar contexts end up with similar vectors. Everything that follows operates on these embeddings, not on the original text directly.
Queries, Keys, Values, and Attention Scores FoundationalKnowledge that endures for decades — core principles
Every token's embedding is turned into three separate vectors by three separate learned transformations: a query (what this token is looking for), a key (what this token offers to others looking), and a value (what this token actually contributes once attended to). A token's query is compared against every other token's key — a simple dot product, higher where the two align — and those comparisons are turned into a set of weights (via softmax, the same normalisation covered on the feedforward networks page) that sum to 1 across the whole sequence. The token's new, contextualised representation is then just the weighted combination of every token's value vector, using those weights.
This is self-attention: every position in the sequence looks directly at every other position in a single step, with no chain of intermediate positions to route through — token 1 and token 1,000 are one attention step apart, not a thousand recurrent steps apart, which is both why transformers handle long-range dependencies more reliably than plain RNNs and why every position's attention can be computed in parallel rather than one step at a time. Real transformers run several of these attention computations side by side — multiple heads — each free to learn a different kind of relationship (one head might track grammatical agreement, another topical similarity), and their outputs are combined afterward.every position looks at every other position, in one step
Positional Information, and What Surrounds Attention FoundationalKnowledge that endures for decades — core principles
Attention alone has no notion of order — swapping two tokens' positions changes nothing about which keys and queries match, which is a problem for language, where word order carries real meaning. Positional encoding fixes this by adding information about each token's position directly into its embedding before attention ever runs. Each transformer layer also includes an ordinary feedforward sublayer (the same building block covered on the feedforward networks page) applied to every position independently after attention, plus residual connections that add each sublayer's input back onto its output, which in practice makes very deep stacks of these layers substantially easier to train.
Related Topics
- Recurrent Neural Networks — the step-by-step alternative this page's attention mechanism replaces, and the specific bottleneck (long chains, sequential computation) it was designed to fix.
- Understanding Large Language Models — this page's mechanics put to work at the scale of an actual language model, covered from a non-technical angle.
- Neural Network Architectures: From Perceptrons to Transformers — the same architecture, placed in its lineage alongside CNNs and RNNs.
References
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017), 5998–6008. https://arxiv.org/abs/1706.03762 ↩