A new breakdown every week — read the newsletter →
PeopleBusinessesTools
Strategies Mental ModelsDecision Tools Business ModelsFrameworksMoats
Learn Book SummariesReading Lists GuidesQuote CollectionsLearning Paths
Artificial IntelligenceNewsletter AboutContact
Foundations

Scaling Laws

The empirical finding that more data, compute, and parameters reliably make models better.

Overview

Scaling laws are the empirical observation that a model's performance improves predictably as you increase three things: the amount of training data, the amount of compute, and the number of model parameters. Crucially, the improvement follows a smooth curve, so labs can forecast how much better a model will get before they build it.

This finding reframed AI research as, in part, an engineering and capital problem: if bigger reliably means better, then whoever can marshal the most data and compute has an edge. It's a reinforcing loop — capability funds revenue, which funds more compute, which funds more capability.

How it works

Step 1

Add data

More high-quality training data improves the model, up to a point set by the other factors.

Step 2

Add compute

More training compute lets the model learn more from that data.

Step 3

Add parameters

Larger models have more capacity to capture patterns — balanced against data and compute.

Step 4

Forecast

Because the curve is smooth, labs predict a bigger model's capability in advance and invest accordingly.

A concrete example

A lab can run small experiments, fit the scaling curve, and forecast that a model 10x larger will hit a particular capability — then raise the capital to build it. That predictability is why frontier AI became a race to assemble the most compute.

Limits & risks

  • Scaling has limits — data and useful compute aren't infinite, and returns eventually bend.
  • Bigger models cost more to train and run, raising energy and capital barriers.
  • Capability gains don't automatically bring safety or reliability.

Frequently asked questions

What are scaling laws in AI?

The empirical pattern that model performance improves smoothly and predictably as you increase training data, compute, and parameters.

Will scaling continue forever?

No — data and economically useful compute are finite, and returns eventually diminish, which is why labs also pursue better algorithms and training methods.

Related