The AI revolution is optimizing for the wrong things


Maz Ahmadi
Image Credits Credit: Maz Ahmadi

TL;DR

Frontier AI models are getting better at coding and agentic tasks but worse at general writing. Wizard Labs founder Maz Ahmadi reports measurable regression on client-specific prose benchmarks across model upgrades. With 44% of organizations scaling AI enterprise-wide but only 37% seeing EBIT impact (McKinsey), the gap between deployment and value is widening. His prescription: build custom evaluation suites before development begins, and stop selecting models based on public benchmarks.

Every large language model existing today is getting worse at writing, and almost nobody is measuring it.

The industry keeps asking which model is smartest, but the bigger question is: smartest at what?

As tech giants chase lucrative enterprise contracts, they are optimizing frontier models for coding, logic, and autonomous agent workflows. In the process, general prose has become an afterthought. On our client-specific benchmarks, my team noticed a regression that should alarm anyone using AI for communication. As models upgrade, their writing performance is actively declining.

That may sound like an esoteric problem for engineers and content teams. It is not. AI now sits inside the daily work of millions of people. Nearly nine in ten companies now use AI regularly in one business function: drafting reports, answering customers, preparing legal documents, writing code, analyzing research, and making decisions. If the underlying models are being optimized for a different definition of performance, everyone inherits the consequences.

The latest debate over AI text watermarking makes this even more interesting.

Anthropic announced this month that future Claude models will watermark generated text to comply with the EU AI Act. The technique uses statistical patterns in token selection so that generated text can later be identified as likely having involved Claude. Anthropic insists its method has no practical effect on quality, creativity, or readability, and points to research behind the approach.

Still, enterprise users should be asking what happens when a model is simultaneously being optimized around regulatory requirements, agentic performance, coding, and reasoning.

I have already seen evidence in our own work that model upgrades can produce worse results for specific writing tasks. Watermarking may prove harmless in isolation; I would not claim otherwise. But the broader trend deserves scrutiny. If frontier labs optimize around the benchmarks that drive adoption and revenue, a model can become objectively better while becoming worse for a particular business.

That is the paradox enterprises are beginning to encounter.

44% of organizations now report AI scaling across the enterprise, up from 38% a year earlier. Yet only 37% report any enterprise-level EBIT impact from AI, while 80% say AI has improved individual productivity.

The technology is spreading faster than the value. I see the same mistake repeatedly. A company asks, “What is the best model?” It selects the model dominating public benchmarks, builds a system around it, and expects the project to behave like conventional software development.

AI does not work that way.

A model that is the best generalized large language model may be mediocre at a company’s particular workflow. The only meaningful test is the task itself. That means developing custom evaluation suites before development begins, measuring the models against the company’s actual requirements, and then continuously testing those results as models change.

This is where enterprise AI needs a new discipline. Start with the business problem, test the technology against real-world performance, and keep refining it until the economics and the workflow make sense, ensuring that the technology supports the way the business actually works before it is scaled across the organization.

A stronger future can exist, where companies build AI around their proprietary data and expertize, giving employees decision-support systems that can extend specialized knowledge across the organization. The danger lies in handing critical judgment to generic systems, treating benchmark scores as proof of competence, and discovering too late that the technology was optimized for someone else’s definition of success.

I believe the first future belongs to companies willing to question the defaults.

That may also mean reconsidering the assumption that every enterprise should rely on the largest proprietary model. Open-weight models are becoming increasingly capable, and they can be hosted within an organization’s chosen infrastructure. The question I hear from enterprises about whether models developed in China are inherently unsafe is often framed as a geopolitical question. Technically, the more important question is where the model runs, who controls the infrastructure, what data leaves the environment, and what security architecture surrounds it.

The competitive advantage will come from making those choices intentionally.

The nature of the AI revolution has shifted from a race for raw intelligence to a discipline of engineering fit. Future market leaders will not win simply by deploying the latest foundational models. Instead, success will favor organizations that define precise operational needs, enforce rigorous benchmarks, and maintain the strategic clarity to replace a hyped, celebrated model when a less glamorous alternative delivers superior results.

Executives should stop asking their AI vendors which model is best. They should ask them to prove which model is best for their business. That single shift would turn AI from a technology procurement exercise into what it actually is: a long-term test of competitive survival.

Get the TNW newsletter

Get the most important tech news in your inbox each week.