OpenAI’s new Ultrafast mode runs GPT-5.6 Sol 14 times faster, on Cerebras chips

The preview is a bet that latency, not just intelligence, is what will make AI agents genuinely usable, and a marquee win for the wafer-scale chipmaker Cerebras.


OpenAI’s new Ultrafast mode runs GPT-5.6 Sol 14 times faster, on Cerebras chips
Image Credits Credit: Shutterstock

OpenAI wants its cleverest model to also be its quickest. The company has previewed Ultrafast, a new tier of its API that runs the flagship GPT-5.6 Sol at up to 14 times the usual speed, reaching around 750 output tokens a second, on hardware built by the wafer-scale chipmaker Cerebras.

Ultrafast is not a new model so much as a new way to serve an existing one. It leans on Cerebras’s outsized chips to strip out the latency that has long dogged frontier AI, and OpenAI opened a limited preview on 13 August to a small group of customers, with plans to widen access as capacity allows.

The pitch turns on a trade-off OpenAI says it can finally dissolve. Until now, anyone who wanted genuinely real-time responses had to drop down to a smaller, less capable model, accepting less intelligence in exchange for speed.

Ultrafast is meant to deliver frontier-grade reasoning and near-instant answers at once, rather than forcing a choice between them.

That combination matters most for the agentic software the whole industry is chasing. An AI agent that has to think for thirty seconds before every step is a demo, whereas one that answers in the time it takes to hold a conversation starts to feel like a product.

Speed, in other words, is quietly becoming a feature as important as raw cleverness.

OpenAI is aiming the tier squarely at time-sensitive work, including incident response and debugging, financial research and fraud detection, real-time customer support and voice, and e-commerce.

Early testers such as Jane Street, Podium, Basis and Rogo describe the change as qualitative rather than incremental, with one saying that speed “completely changes the call experience for complex work” and unlocks “synchronous experiences for users that were previously limited by intelligence.”

For Cerebras, the deal is a marquee endorsement at a helpful moment. The company went public in one of the year’s biggest listings but has since struggled to convince the market that wafer-scale ambition translates into durable profit, so powering OpenAI’s fastest tier is precisely the kind of validation it needed.

It is also a reminder that the exotic chip architectures once dismissed as science projects are now doing real work for the biggest names in AI.

The speed itself comes from an unusual piece of engineering. Cerebras builds processors the size of a dinner plate, cut from a single silicon wafer, which lets an entire model sit on one chip rather than being split across racks of Nvidia GPUs that must constantly shuttle data between them.

Stripping out that internal traffic is what collapses the delay between a prompt and a reply, and it is the basis of the company’s long-running argument that its design suits inference far better than the general-purpose chips built for training.

The move fits a broader shift, too. As raw model capability begins to plateau, the contest is moving toward who can run those models fastest and most cheaply, a race that has lifted inference specialists like Groq and turned latency into a selling point.

Rivals such as SambaNova and a clutch of specialist clouds are chasing the same prize, and the market increasingly rewards whoever can make a given model answer soonest, not simply whoever trained the biggest one.

For OpenAI, leaning on Cerebras is also a quiet step away from total dependence on Nvidia, of a piece with its work on its own custom silicon.

OpenAI has not published pricing, and running its top model at 14 times the speed on specialist hardware is unlikely to come cheap, so Ultrafast may remain a premium option for latency-obsessed cases rather than a default setting.

It is also just a preview, gated to a handful of customers while OpenAI hunts for capacity. Even so, the message is plain enough. In the next phase of the AI race, being clever will not count for much if you are also slow.

Get the TNW newsletter

Get the most important tech news in your inbox each week.