Speed is no longer a selling point for AI agents, caution is


AI Agents graphics
Image Credits Credit: Canva

AI agents have spent two years competing on the same thing: independence. Less oversight, fewer clicks, more autonomy, that was the pitch. It didn’t matter if the agent lived in a browser extension, a cowork app, or a coding tool. I think that pitch has become the industry’s biggest liability. At Decodo, we recently reviewed dozens of these agents and tested them ourselves. We are at the point where businesses are starting to rely on the agents that know how to switch caution on when the moment calls for it.

Everyone can already do the easy part

Two years ago, chaining a handful of actions into one task was a genuine differentiator. “Save hours of manual work,” every landing page said. It’s not a winning point anymore. Most agents on the market, browser-based or not, can now string several steps together. They plug into outside tools without much trouble. That race is over, and I don’t think it was ever the hard part. Capability parity is what happens once enough vendors copy a feature. This category hit that point faster than most.

Money is where the real test starts. A growing number of agents can now complete a purchase outright, on their own. They no longer just queue one up for a human to approve. That’s a genuine achievement, one that touches payment details, order accuracy, and liability all at once. But completing a purchase and knowing when not to are two different skills, and almost nobody has built both.

In our own research, only three agents documented both a working checkout and a clear pause first. That’s the moment an agent checks with the user before it spends anything. All three came from Amazon or OpenAI. Building a pause button is a policy decision as much as an engineering one. Right now, only the companies with the most legal and reputational exposure seem willing to make it.

The industry keeps relearning the same lesson

None of this is new, it’s just moving faster than usual. Online payments went through the identical arc, convenience shipped first, friction got added only after fraud losses forced the issue. Card-not-present fraud existed for years before anyone added a one-time password to a checkout flow. I think agentic commerce is running the same playbook on a much shorter timeline. I don’t expect a regulator to be the one who forces the correction.

It’ll more likely be a procurement or security team. Enterprise buyers already run vendors through security reviews before granting access to anything sensitive. Once an agent with no confirmation step causes one expensive, public mistake, and I think that’s a matter of when, those reviews will start asking for a pause button by name. That’s a faster, blunter forcing function than legislation, and it doesn’t need a single regulator to agree on anything first.

What we keep seeing when we test these tools ourselves

A vendor’s public documentation is one layer of evidence, and it’s the layer every buyer sees first. It usually isn’t the whole picture.

At Decodo, we ran hands-on tests against a handful of the boldest claims in the market. Most of them didn’t hold up. The agent with the single best documentation score in our review still got stuck on complex layouts. It couldn’t clear a simple CAPTCHA once we tried it.

That specific failure mode is one I recognize. We spend our time at Decodo solving web access at scale, CAPTCHAs, anti-bot walls, and layouts that shift depending on who’s asking. Most of the industry talks about agent failures as reasoning problems, a model that didn’t plan well enough. In my experience, a good number of them are access problems. The agent never reliably saw the page it was supposed to act on. And that’s a plumbing failure.

Another agent’s own promotional demo showed it fixing bugs and files that never existed in the referenced code. That’s a strong sign the demo was staged. When we tested that same agent on 10 real tasks, it completed two. A documented score is a floor a vendor sets for itself.

None of this means agentic tools should be avoided. It means the buying criteria haven’t caught up with the marketing. What matters now goes beyond what a task list can show you. Before evaluating a task list, I’d check a couple of things:

  • Confirmation policy before irreversible actions
  • What happens when a page contains hidden instructions or fails to load properly.

Checkout is mostly solved now, at least on paper. The harder problem, knowing when to stop, is a choice, and most agents are still avoiding it. I think that avoidance gets more expensive every quarter this category keeps shipping without a pause button. Whoever builds restraint in first will define what “trustworthy” means for everyone else.

Get the TNW newsletter

Get the most important tech news in your inbox each week.