Nvidia's GPUs built the AI boom we're living through. They're still the go-to choice for training the biggest AI models. But something is changing fast underneath that. Google, Amazon, Microsoft, and Meta have all built their own AI chips. OpenAI is doing the same. A few startups think GPUs were never really the right shape for AI. After all, GPUs were built for video games first. None of this has knocked Nvidia off its throne yet. It still controls about 70 to 80% of the AI chip market. But the fastest-growing part of that market is no longer GPUs.
Training vs. inference: what's really driving this
To understand why, you need to know there are two different jobs in AI. Training is the process of teaching a model. It happens once, and it's expensive. GPUs are great at training because they can flex across many kinds of math. Inference is different. That's what happens every time someone actually uses the model, like asking it a question or making an image. Inference now makes up about two-thirds of all AI computing. In 2026, big cloud companies spent more on inference than on training for the first time. Inference does the same kind of work over and over. That makes it a better fit for a chip built to do one job well, instead of a GPU built to do everything.
Big tech companies are already doing this
This isn't just talk. It's already happening. Google's newest chip, called Ironwood, came out in late 2025. Anthropic says it now uses more than a million Ironwood chips to run Claude. That's the first time a custom chip has hit that scale with one customer. Amazon has deployed over 500,000 of its own Trainium2 chips. Anthropic trains models on them inside one of Amazon's biggest data centers. Microsoft has a chip called Maia. Meta has its own chip family called MTIA, built mainly for inference. All of these get design help from Broadcom or Marvell, two companies that have quietly become essential to this shift. And they're all made by TSMC, the same factory that builds Nvidia's chips too.
OpenAI doesn't have its own hardware yet, but it's moving fast. It has a roughly $10 billion deal with Broadcom to build custom chips. The goal is 10 gigawatts of computing power by 2029. It also has a similar deal with Cerebras that runs through 2028. Analysts expect custom AI chips to grow about 44% in 2026. Regular GPUs are only expected to grow about 16%. That doesn't mean GPUs are shrinking. Both are still growing. Custom chips are just growing faster from a smaller starting point.
Some companies are rebuilding the chip from scratch
Some companies aren't just tweaking the GPU idea. They're throwing it out completely. Cerebras builds something called a wafer-scale chip. Normally, a silicon wafer gets cut into lots of small chips. Cerebras uses almost the whole wafer as one giant chip instead. It packs in about four trillion transistors. This design skips the delays that happen when data has to travel between separate chips. Cerebras says this makes it faster for inference. OpenAI backed that claim with a $10 billion bet on the company. That's a real sign that at least one big customer believes it, even though it hasn't been independently proven.
Groq took a different path. It built a chip with a fully predictable, step-by-step design, instead of a GPU's more flexible one. In late 2025, Nvidia made a roughly $20 billion deal for Groq's technology. It's worth getting this right: it wasn't a full buyout. It was a non-exclusive license. Groq's founder and top leaders moved over to Nvidia to help scale the tech. Groq itself keeps running as its own company, just without its cloud business, which wasn't part of the deal. Either way, it's the biggest deal Nvidia has ever made. It shows Nvidia would rather absorb a threat than fight it head-on. Further out, a company called IonQ is exploring whether quantum computers could help with small parts of AI work, mostly optimization problems. But even IonQ says quantum computing will support regular AI chips, not replace them.
Why this matters, and what's still uncertain
The practical reason this matters comes down to cost and power. A chip built for one specific job can run cheaper and use less energy than a general-purpose GPU doing the same task. At the scale big tech companies operate, that difference adds up to billions of dollars. It also puts real pressure on power grids. There's a strategic reason too. A company that builds its own chip depends less on Nvidia's pricing and supply. It also depends less on Nvidia's CUDA software, which has locked developers into Nvidia's world for years.
What's still unclear is how far this goes. Analysts expect custom chips to make up somewhere between 10% and 20% of the training and inference market by the end of 2026. That's up from under 5% just a couple of years ago. But that's a prediction, not something that's already happened. Building a competitive AI chip also costs a huge amount up front, tens of millions of dollars at minimum. That's why this shift is mostly happening at giant companies with deep pockets and huge AI workloads. Smaller AI startups still rely on Nvidia almost by default.