A Year of Testing Has Cooled Security Teams' Enthusiasm for AI-Run Hacking Drills
Companies that hoped AI could fully replace human security testers have pulled back sharply. New data shows only 9% still trust fully automated systems, down from nearly a third just twelve months ago.

Key points
- The share of organisations willing to rely entirely on AI-powered penetration testing, structured drills where security experts try to break into their own systems to find weaknesses before criminals do, fell from 29% in 2025 to just 9% in 2026, according to a June 25 report by Cobalt, a security-testing firm.
- 78% of companies reported that automated security tools missed serious vulnerabilities during assessments, a problem known as false negatives.
- Microsoft patched 206 unique software flaws in its June 2026 monthly security update, a record number, with AI tools credited for finding many of them.
- New vulnerabilities are being reported at a rate 46% higher than analysts predicted, according to a June 2026 analysis by the Forum of Incident Response and Security Teams (FIRST).
- HackerOne, a firm that runs bug-bounty programmes where independent researchers earn money for finding security flaws, paused its Internet Bug Bounty programme because AI-generated submissions overwhelmed its validation staff.
One year ago, roughly three in ten security professionals believed AI could handle their organisation's penetration testing entirely on its own. Pen testing is when trained experts deliberately try to break into a company's own systems, mimicking a real criminal, so weaknesses can be fixed before anyone malicious finds them.
That confidence has collapsed.
A report published June 25 by Cobalt, a penetration-testing-as-a-service firm, found the share of organisations prepared to trust fully autonomous AI systems with this job dropped from 29% to 9% in a single year. Most now want a human expert involved at every meaningful step.
Why did confidence drop so fast?
AI security tools kept missing things that mattered. Nearly four in five companies, 78%, per Cobalt, found that automated tools produced false negatives: the software examined a real, serious flaw and reported everything was fine. A missed vulnerability is an open door.
The tools also generated enormous volumes of output, burying security teams in data. "A human expert is needed to decide whether a lead is worth pursuing," said Derek Rush, managing senior consultant at offensive-security firm Bishop Fox. Without that judgement, teams chase noise while genuine risks go unaddressed.
Costs compounded the frustration. AI-powered security services bill by usage, and those bills have proven difficult to predict, a pattern security leaders have watched play out elsewhere in their organisations.
Gunter Ollmann, chief technology officer at Cobalt, said CISOs, the executives responsible for an organisation's overall security posture, spent the past two years under board pressure to adopt AI everywhere. Now, with a year of real-world results in hand, many have grown cautious. Our 2 July piece on the enduring value of hands-on pen testing showed the judgement gap isn't new; what's new is that organisations are paying for it at scale.
Where does the vulnerability surge fit in?
The problem is partly self-inflicted. AI-assisted programmers are writing more code, faster, and more code means more places for flaws to hide. FIRST analysts calculated that new vulnerabilities are being reported at a rate 46% above last year's forecasts. Microsoft's record June 2026 update, which patched 206 separate flaws, shows exactly where that trend leads. AI found many of them; human engineers then had to verify each one.
That verification step is now the chokepoint.
"The constraint is no longer discovery; it is the human capacity to verify, coordinate, and patch," FIRST analysts Jerry Gamblin and Eireann Leverett wrote. We reported in June how CISOs were already quietly reallocating budgets to handle exactly this kind of pressure.
Should you be concerned about the services you use?
Yes, and here's why it's not theoretical. The organisations holding your data are under growing pressure to find and fix security gaps faster than ever, while their primary detection tools are generating more noise than their teams can process. The durable model, as HackerOne frames it, is AI doing the relentless broad first pass and humans handling depth and judgement. Full autonomy remains out of reach for now, though AI models will keep improving. What to watch: whether verification tooling catches up before the gap between discovery and patching becomes a reliable criminal opportunity.



