AI Learning and Artificial Intelligence
OWASP’s 2026 Top 10 for LLM applications moved excessive agency from sixth to third and dropped improper output handling from fifth to tenth, in the first edition weighted partly on incident data. Europe’s Cyber Resilience Act began requiring 24-hour vulnerability reports on 11 September, but it governs products rather than how an agent is deployed.
Excessive agency has moved from sixth place to third in OWASP’s 2026 top ten for LLM applications, and improper output handling has fallen from fifth to last. The hardest problems now sit outside the model, Steve Wilson argues.
Excessive agency is not about what a model says. It is about what it has been handed: the tools, credentials and scopes it can reach once it answers.
The 2026 edition was the first weighted partly on recorded incidents, 6,639 of them, counting for a quarter of the ranking against three quarters practitioner consensus.
Wilson co-led the project with Rock Lambros and is chief AI officer at Exabeam, a security company. Read the column as both things at once.
His prescription is unglamorous. Read-only tools instead of general connectors, requests inside the user’s own scoped identity, a policy enforcement point between the model and anything downstream, and human approval for anything hard to reverse. It echoes the gap a founder described to us in August, between what an agent can do and what it should be authorised to do.
Europe started a clock five days ago.
Since 11 September, manufacturers have had 24 hours to warn their national response team of an actively exploited vulnerability, under the Cyber Resilience Act. A fuller notification follows at 72 hours, through a platform ENISA switched on the same day.
Penalties reach EUR 15M or 2.5% of worldwide turnover. The rest of the regulation applies from 11 December 2027.
Our own reporting found the bottleneck is not obtaining a bill of materials but trusting one. A list is worth only the confidence you can place in it.
The mismatch is what the clock points at. The Act governs products with digital elements, and lawyers reading it place cloud-delivered software outside, which is how most agents reach a customer.
Four teams broke AI agents four ways in ten days this summer, and the same flaw ran through all of them. The failures sat in the ecosystem around the model, not the model.
Wilson’s own figures make the point. On his internal red team suite the oldest model fails 17% of the time and the newest 2%, and he says 98% is nowhere near good enough when every failure leaks data.
The one European report anybody has confirmed came under the AI Act, not this one. OpenAI filed an incident report over agents that occupied a German wiki. Agency is falling between two regimes.
Get the TNW newsletter
Get the most important tech news in your inbox each week.