A startup says the AI bottleneck isn’t compute. It’s memory, and it ditched the GPU to prove it.

Majestic Labs has unveiled a server that drops the GPU for Arm cores and up to 128TB of cheap LPDDR6 memory. It claims to shame a rack of Nvidia chips on memory and power. None of it has shipped or been independently tested.


A startup says the AI bottleneck isn’t compute. It’s memory, and it ditched the GPU to prove it.

The AI hardware conversation is stuck on one word: compute. A startup out of Tel Aviv wants to change the word to memory. Majestic Labs, founded in 2023 by former Google and Meta engineers, has unveiled a server it says can do the work of a rack of Nvidia GPUs. It does so by attacking a different bottleneck.

The pitch, reported by TechRadar, is that pairing pricey GPUs with scarce high-bandwidth memory has become a dead end for AI inference. Running a model is often limited by how much fast memory you can reach, not raw compute. So Majestic ditched the GPU.

Its server, Prometheus, swaps graphics chips for what it calls Ignite AI Processing Units. Each unit blends Arm cores with RISC-V vector and tensor engines. Up to 12 sit in one server, sharing a single pool of 8TB to 128TB of LPDDR6. That is the cheap memory found in phones, not the costly high-bandwidth memory that GPUs depend on.

A different way to hit the memory wall

The trick is the wiring. Instead of bolting memory onto each GPU package, Majestic pools it through custom aggregation chiplets. Copper cables up to a metre long connect them. The result, it claims, is one coherent pool far larger than a GPU box can address.

The 💜 of EU tech

The latest rumblings from the EU tech scene, a story from our wise ol' founder Boris, and some questionable AI art. It's free, every week, in your inbox. Sign up now!

The comparison it reaches for is stark. An Nvidia DGX B300 with eight Blackwell GPUs carries 2.3TB of high-bandwidth memory. Majestic says Prometheus offers more than 50 times as much fast memory, at 1.7 times the interconnect bandwidth. One rack, it claims, matches 25 of Nvidia’s Vera Rubin racks for fast memory, at a fraction of the power.

The software story is friendlier than the hardware sounds. Prometheus is built to open standards and supports PyTorch, vLLM and OpenAI’s Triton. Models built for GPUs are meant to run on it without changes. Majestic says it has already taken orders from large enterprises, neoclouds and hyperscalers.

Striking numbers, no shipped hardware

Every figure here carries the same asterisk. It is Majestic’s own, ahead of independent testing, and nothing has shipped. The company has about 40 staff across Tel Aviv and Los Angeles. It raised $100m late last year, a modest sum against what its rivals command.

There are physical questions too. A 128TB pool built from 2GB LPDDR6 dies would need roughly 64,000 of them. That implies more than a hundred aggregation chiplets in a single server. It is a lot of parts to keep coherent, and TechRadar notes buyers will likely wait for independent benchmarks before switching.

What Majestic joins is a growing line of startups attacking Nvidia from odd angles. Optical chips, edge silicon for inference and open networking gear each pick a different weak point. The common thread is that Nvidia’s dominance is now a target from several sides at once, even its software moat.

The memory-wall pitch is the newest of them. It is also the least proven. Majestic has drawn a picture of a rack that shames a room of GPUs on memory and power. Next year, when hardware ships and someone else runs the benchmarks, the picture either holds or it does not.

Get the TNW newsletter

Get the most important tech news in your inbox each week.