In 2023, Falcon 180B was one of the biggest AI language models ever released to the public — 180 billion parameters, trained on huge amounts of text. A year later, Meta's Llama 3 8B, a model about 22 times smaller, started beating it on standard tests.
That's not just a curiosity. It's a sign of a bigger shift: simply making AI models bigger has stopped being the best way to make them better. For most of what businesses actually need AI to do, a smaller, cheaper model often works just as well — or better.
Why "bigger is better" became the default
For years, AI progress followed a simple pattern: make a model bigger, feed it more data, give it more computing power, and it gets better in a fairly predictable way. This idea, known as scaling laws, shaped how the whole industry invested, producing the huge, expensive models we now recognise, like GPT-4 and Gemini.
But that focus on size quietly ignored a more practical question: once a model is built, what does it cost to actually run it, day after day, for years?
The real cost is running it, not building it
A model is trained once. But it answers questions millions or billions of times afterward. Industry studies show that running a model (called inference) makes up 80 to 90% of its total lifetime cost, while training is often just 10 to 20%. A model that's slightly better but far more expensive to run every single time can lose that trade-off badly once it's actually being used at scale.
This is the part often missed: companies aren't shrinking models because bigger models got worse. They're doing it because the bill for running a model never stops arriving — and that ongoing cost is what decides whether an AI product actually makes money.
Bigger models also hit diminishing returns
As models are fed more and more data, each new piece of data teaches them a little less than the last. Researcher Sara Hooker documented this in a widely discussed 2026 essay, using the Falcon-versus-Llama example above to show how smaller models are increasingly catching up to bigger ones.
Most everyday business tasks — sorting a support ticket, pulling details from an invoice, summarising a document — don't need the broad reasoning ability that justified building a huge model in the first place. Paying to run a massive model for a job a small one can handle is simply wasted cost. That's the core idea behind the shift: match the model's size to what the task actually needs.
How companies actually shrink a model
Two methods do most of the work. Distillation trains a smaller “student” model to copy the behaviour of a larger “teacher” model, learning from its answers instead of starting from scratch. Google confirms its smaller Gemini models are built this way specifically to cut running costs.
Quantisation shrinks a model by storing its internal numbers with less precision — similar to compressing a high-resolution photo. Most of the detail stays intact, but the file becomes much smaller and faster to run. Combined with open, freely downloadable models from companies like Google, Meta and Alibaba, these techniques have made small, efficient models available to almost anyone, not just big AI labs.
Small, but not a replacement for everything
The numbers back this up: a 2026 MIT Sloan study found open models cost roughly 87% less to run than closed ones, while reaching about 90% of their performance within months of release. Smaller models also run directly on phones and laptops, keeping data private and working even offline.
Still, small models generally can't match large ones on open-ended reasoning or unfamiliar problems. Most companies aren't replacing big models — they're using small ones for routine work and saving expensive, powerful models for harder tasks. Choosing the right size for the job, instead of always reaching for the biggest model, is now the smarter default.