Nvidia’s real moat was never the chips. AI has started rewriting it.

Nvidia’s grip on AI rests on CUDA, the software layer around its chips, not the chips alone. AI coding agents and the shift to inference are the first credible threats to it. They may also hand Nvidia its next moat.


Nvidia’s real moat was never the chips. AI has started rewriting it.

The usual story about Nvidia is a hardware story: the fastest chips, in the shortest supply, at the highest price. The more important part is software. For two decades the company’s real moat has been CUDA. It is the layer that turns its silicon into something developers can build AI on. That moat is now being tested.

CUDA, short for Compute Unified Device Architecture, took years to build. It bundles ready-made code and debugging tools. It also lets thousands of chips train a model together. The new threat is blunt: AI coding agents that can write that kind of low-level software themselves.

Jeremy Nixon is a former Google Brain researcher who founded the startup Infinity. He told Business Insider his team used agents to rebuild CUDA-like software for chip firm D-Matrix in about 10 hours. He framed it as proof that one of Nvidia’s biggest moats is being crossed.

Two moats, both under pressure

CUDA’s first advantage is the software. Its second is everything built on top of it. Millions of lines of company code and workflows make switching to a rival chip slow and costly. Amazon’s own documents once flagged CUDA as a major roadblock to adopting its in-house AI chips.

The 💜 of EU tech

The latest rumblings from the EU tech scene, a story from our wise ol' founder Boris, and some questionable AI art. It's free, every week, in your inbox. Sign up now!

Agents chip away at the first advantage. The pressure is not only from startups. Google, Amazon and Microsoft have spent years writing software for their own chips. OpenAI and Anthropic have shown models that can generate system code. DeepSeek’s founder said coding agents, plus its own programming language, made building AI software much easier.

Nvidia does not dispute the trend so much as claim it. It says developers lean on CUDA’s libraries more every year. It also uses coding agents to build CUDA faster itself. The lock-in, in other words, may bend before it breaks.

The inference problem

The sharper threat is a shift in what AI chips are for. As the industry moves from training models to running them, priorities change. Buyers care less about peak performance and more about running AI cheaply. That favours software which works across different chips, not software welded to one vendor.

On the inference side, CUDA “is no longer a factor,” said Marshall Choy of Korean chip startup Rebellions. He calls it an open source play. It is the same opening our other inference challengers are chasing. Think optical chips, networking silicon, and Alibaba’s open-source alternative to CUDA itself.

Chris Lattner, whose startup Modular builds chip-agnostic AI software, says CUDA’s age cuts both ways. It carries years of legacy from its gaming origins, he says, “like Microsoft Windows trying to fit onto a phone.” Wall Street has noticed. Analysts read Nvidia’s flat stock over the past year as partly a bet against the moat.

The moat may just move

The counter-case is that agents relocate the moat rather than remove it. Generated code still has to be verified and optimised, and that is where CUDA’s ecosystem is deepest. Bing Xu, whose last chip-software startup was bought by Nvidia, argues that verification becomes the next moat.

“Agents can generate a lot of code in a short time, but verification is the biggest bottleneck,” Xu said. Lattner is blunter still. The hype, he says, is “very overblown.” Writing code is a small part of building software; the hard part is tuning it for production. And chip software is a niche field, with few public examples for agents to learn from.

So the honest read is not that CUDA is falling. It is that the thing protecting Nvidia has, for the first time, a plausible expiry date. Agents can now do in hours what used to take specialist teams years. Inference rewards whoever frees buyers from a single stack. Whether the moat holds comes down to one thing. Rivals must close the gap faster than Nvidia, which is “not sleeping,” can open a new one.

Get the TNW newsletter

Get the most important tech news in your inbox each week.