AMD announced the deal on Thursday, after the market closed, and did not say what it paid. Taalas, founded in 2023, builds what it calls model-specific chips. Instead of loading weights from memory, it bakes them straight into the silicon. AMD says the technology will join its accelerator roadmap, working alongside its Instinct GPUs.
“I’m a big believer that there’s no one-size-fits-all as it comes to chips,” AMD chief executive Lisa Su said at a July event. That line is the whole logic of the deal. The industry spent four years buying general-purpose GPUs. AMD is betting the next phase rewards something far narrower.
What Taalas actually built
A normal AI chip keeps memory and compute apart, then burns huge effort shuttling data between them. That gap is why modern systems need stacked memory, exotic packaging, and liquid cooling. Taalas merges the two. Its co-founder, Ljubisa Bajic, says the result needs none of it: no high-bandwidth memory, no 3D stacking, no liquid cooling.
The speed claims are startling. Taalas’s first chip runs Meta’s Llama 3.1 model at about 17,000 tokens a second per user, which it says is many times faster than a leading GPU, at a fraction of the cost and power. Those are the company’s own numbers. The first version also leans on aggressive compression that dents output quality.
You had better love that model
Here is the trade. Hardwire a model into a chip and you marry it. New models now arrive almost monthly, and a serious change means re-spinning the silicon. Taalas softens the blow: only two of the chip’s metal layers need redoing, so a new model can be etched in about two months rather than six.
Even so, buyers will have to be very sure of their pick. “You better really love that model,” as The Register put it. In practice the customers are likely to be the big model labs and inference providers, the few players confident enough about a model to cast it in silicon.
The bigger race
AMD is not alone. The deal echoes Nvidia’s $20bn Groq purchase in December, and Qualcomm’s acquisition of the compiler startup Modular in July. Anthropic is building its own silicon team to shape hardware around its models. Everyone is converging on one idea: match the chip to the workload.
For AMD, Taalas slots into a bigger push. It has spent the year assembling an inference stack around its Helios racks and signing gigawatt-scale deals with the likes of Anthropic and OpenAI. Those sell flexible GPUs by the gigawatt. Taalas offers the opposite: extreme efficiency for a model that has stopped moving.
The purpose underneath is a challenge to Nvidia, whose general-purpose chips made it the world’s most valuable company. As AI shifts from training models to serving them, that grip loosens a little. AMD has bet that the future of inference is narrow, fast, and etched in place. The deal is expected to close in the fourth quarter.
Get the TNW newsletter
Get the most important tech news in your inbox each week.