Writer bets on cheaper AI agents with Palmyra X6 and a leaner harness

The enterprise AI firm says its new flagship model and a smarter orchestration layer can cut token costs by up to half, and it is not shy about blaming the big labs for the bills.


Writer bets on cheaper AI agents with Palmyra X6 and a leaner harness
Image Credits Credit: Writer

Writer, the enterprise AI company, has launched Palmyra X6, its new flagship model, alongside an upgraded “harness” built to do something the industry has been curiously reluctant to promise: spend fewer tokens.

The model is a post-training variation on Z.ai’s open-source GLM-5.2, and it is available to Writer’s clients from today.

The pitch arrives at a nervy moment. As agentic AI quietly multiplies the steps, and therefore the tokens, behind every task, enterprise bills have started to sting.

Writer is wagering that a cost backlash has arrived, and it is far from alone in scenting the opportunity, with Baseten raising $1.5bn on the theory that AI’s profits lie in cheap inference.

A harness, in Writer’s telling, is the orchestration layer wrapped around a model, the machinery that decides how a multi-step agent actually executes each request.

Optimise that layer, trimming the redundant calls and bloated context that agents tend to accumulate, and you cut token consumption without ever touching the model itself.

The savings, Writer says, are substantial. The company estimates that the new model and harness together cut costs by as much as 50% for basic tasks, while a Writer research paper found that harness-efficiency changes alone trimmed costs by roughly 40% on average across testing.

What makes the harness interesting is that it is model-agnostic. It works with Writer’s own models, but also with external ones served through Microsoft Azure and Amazon Bedrock, so the efficiency gains are not locked to a single vendor’s roadmap.

“The harness is the one component whose efficiency multiplies across every model an organization runs, present and future,” Writer’s researchers write, framing orchestration rather than raw model quality as the lever enterprises have been ignoring.

It is also a pointed one. Writer’s underlying claim is that the big labs have little incentive to help you spend less, because their revenue rises with every token you burn. Distrust of that arrangement is doing a lot of work in the company’s messaging.

CEO May Habib put it bluntly. “The enterprise is absolutely sick of chasing the next benchmark,” she said. “They want flattening cost.” It is a neat inversion of the usual sales script, in which each new model is faster, cleverer and, invariably, hungrier.

A wider thrift-maxxing trend, in which buyers reach for cheaper Chinese models, has already started to unsettle the valuations underpinning eventual OpenAI and Anthropic IPOs, a sign that price sensitivity is no longer a fringe concern.

Building the flagship on GLM-5.2 is telling in its own right. Rather than train a frontier model from scratch, Writer took a capable open-source base and specialised it through post-training, a route that keeps costs down and neatly matches the frugal story it is selling.

For European buyers, the framing should feel familiar. Scepticism about American hyperscalers’ incentives, and a preference for architectures that are not hostage to one provider, has been a running theme on this side of the Atlantic, and Writer is pitching straight into it.

There are caveats worth keeping in mind. Efficiency figures produced by a vendor about its own product deserve the usual pinch of salt, and “up to” 50% is not the same as a guaranteed 50%, but the direction of travel is hard to dispute.

The deeper shift is what Writer calls harness engineering: the discipline of spending fewer tokens per task rather than chasing another point on a leaderboard. If it takes hold, and the incentives increasingly suggest it will, the era of the ever-hungrier model may finally be meeting its budget.

Get the TNW newsletter

Get the most important tech news in your inbox each week.