A new breakdown every week — read the newsletter →
PeopleBusinessesTools
Strategies Mental ModelsDecision Tools Business ModelsFrameworksMoats
Learn Book SummariesReading Lists GuidesQuote CollectionsLearning Paths
Artificial IntelligenceNewsletter AboutContact
Foundations

The Transformer

The 2017 architecture, built on 'attention', that made modern AI possible.

Overview

The Transformer is the neural-network architecture, introduced in 2017, that underlies essentially every modern large language model. Its breakthrough was a mechanism called self-attention, which lets the model weigh the relevance of every word to every other word in a sequence, in parallel, rather than reading strictly left to right.

That parallelism is why Transformers scale so well on modern hardware — and scale, it turned out, was the thing that unlocked the capabilities we now associate with AI.

How it works

Step 1

Tokenise

Text is broken into tokens and turned into numeric vectors the model can process.

Step 2

Attention

Self-attention lets each token 'look at' every other token and weigh what matters for meaning.

Step 3

Layers

Stacks of attention and processing layers build up increasingly abstract representations.

Step 4

Predict

The final layer outputs a probability for the next token, and the process repeats.

A concrete example

In the sentence 'the trophy didn't fit in the case because it was too big', attention lets the model connect 'it' to 'trophy' rather than 'case' — resolving the ambiguity by weighing the relationships between all the words at once, which older architectures struggled to do.

Limits & risks

  • Attention cost grows with sequence length, making very long context expensive.
  • Transformers need enormous data and compute to reach their potential.
  • The architecture is powerful but opaque — hard to interpret what it has learned.

Frequently asked questions

What is 'attention'?

A mechanism that lets the model weigh how relevant every token is to every other token, capturing context and relationships across a whole sequence at once.

Why did the Transformer change everything?

Its parallel design scaled efficiently on modern hardware, and scaling it up produced the leap in capability behind today's AI.

Related