Overview
The Transformer is the neural-network architecture, introduced in 2017, that underlies essentially every modern large language model. Its breakthrough was a mechanism called self-attention, which lets the model weigh the relevance of every word to every other word in a sequence, in parallel, rather than reading strictly left to right.
That parallelism is why Transformers scale so well on modern hardware — and scale, it turned out, was the thing that unlocked the capabilities we now associate with AI.
How it works
Tokenise
Text is broken into tokens and turned into numeric vectors the model can process.
Attention
Self-attention lets each token 'look at' every other token and weigh what matters for meaning.
Layers
Stacks of attention and processing layers build up increasingly abstract representations.
Predict
The final layer outputs a probability for the next token, and the process repeats.
A concrete example
In the sentence 'the trophy didn't fit in the case because it was too big', attention lets the model connect 'it' to 'trophy' rather than 'case' — resolving the ambiguity by weighing the relationships between all the words at once, which older architectures struggled to do.
Limits & risks
- Attention cost grows with sequence length, making very long context expensive.
- Transformers need enormous data and compute to reach their potential.
- The architecture is powerful but opaque — hard to interpret what it has learned.
Frequently asked questions
What is 'attention'?
A mechanism that lets the model weigh how relevant every token is to every other token, capturing context and relationships across a whole sequence at once.
Why did the Transformer change everything?
Its parallel design scaled efficiently on modern hardware, and scaling it up produced the leap in capability behind today's AI.