Cerebras launches the CS-4, its first multi-wafer system, though the chip inside is not new


Cerebras launches the CS-4, its first multi-wafer system, though the chip inside is not new

Cerebras has put three of its dinner-plate-sized processors into a single rack for the first time. The CS-4, unveiled on Tuesday at the company’s Supernova event and shipping this quarter, is pitched as an inference machine for frontier models, and Cerebras says it runs them up to 30 times faster than GPU-based systems.

The launch lands five days after OpenAI’s Ultrafast mode, which runs GPT-5.6 Sol roughly 14 times faster on Cerebras silicon, went live. It is also the first hardware the company has shipped since its $5.55bn Nasdaq debut in May, the largest US tech listing since Snowflake.

On paper the system is formidable. Each CS-4 carries three WSE-3 Turbo wafers for a combined 750 petaflops of sparse FP16 compute, 129.6 petabytes per second of memory bandwidth, and support for models above 50 trillion parameters.

Wafer-to-wafer latency falls to two microseconds from five, and the rack uses half as many components as its predecessor. Cerebras has moved power conversion, in its own phrasing, a hundred times closer to the processors, mounting it in a removable backpack at the rear of the chassis.

What is less clear is whether the chip inside is actually new. The WSE-3 Turbo carries the same four trillion transistors, the same 900,000 cores, the same 44GB of on-chip SRAM, and the same TSMC 5nm node as the WSE-3 it replaces.

The Register concluded it is not new silicon so much as the existing die pushed from roughly 1.4GHz to 2.8GHz. Per-wafer compute and bandwidth both exactly double, which is the signature of a clock bump rather than a redesign, and Cerebras has a genuinely new generation scheduled for 2027 anyway.

The 30-times claim measures tokens per second per user on a single model, gpt-oss-120b, against unnamed GPU systems, and the 250 petaflops per wafer is a sparse FP16 number against 25 petaflops dense.

Where the design does look genuinely differentiated is power. The Register puts the CS-4 at an estimated 120 to 140 kilowatts per rack, roughly half what comparable AMD and Nvidia rack systems draw, which gives the claimed tenfold gain in throughput per watt somewhere plausible to start from.

Andrew Feldman, the chief executive, framed the launch around latency rather than raw throughput. “In AI, speed is productivity,” he said in the announcement, and he told Reuters the company expects to get “four times as fast between now and the end of 2027, and 20 times more throughput”.

Sean Lie, the chief technology officer, tied the speed argument to agents rather than chatbots. Being 30 times faster, he said, “gives an agentic system room for more than an order of magnitude as much reasoning, verification, or tool use”, which is the part of the pitch aimed at enterprise buyers rather than benchmark tables.

Cerebras named OpenAI, G42, MBZUAI, and AWS alongside the launch, though it disclosed no CS-4 customer agreements and no pricing.

AMD’s Helios rack, unveiled in July, sits in the partner list, which is consistent with a company that has said it will work with everyone in AI hardware except Nvidia.

The financial backdrop is mixed. Second-quarter revenue came in at $180.1m, up 74% year on year with cloud revenue nearly quadrupling, but it fell sequentially from $193.4m in the first quarter, and the margin pressure Cerebras flagged in June has not lifted.

The quarter produced a GAAP net loss of $450.5m, an adjusted loss of $6.9m, and $25.4bn in remaining performance obligations. Cerebras guides to $880m to $890m of full-year revenue and 600 megawatts of data centre capacity live and under contract by the end of 2027.

Concentration remains the structural question. G42 and the Mohamed bin Zayed University of Artificial Intelligence together accounted for around 86% of 2025 revenue, and the OpenAI contract signed in January, worth more than $10bn at signature, is the company’s answer to that, alongside second-quarter additions including Cognition, Lovable, CrowdStrike, Block, and Figma.

First CS-4 shipments are due before the end of the quarter. The harder test comes in 2027, when the next generation will have to arrive on new silicon rather than a faster clock.

Get the TNW newsletter

Get the most important tech news in your inbox each week.