Microsoft's AI bug-hunters logged 140 Windows CVEs in four months. The queue is now the problem.

FORGE Lab's agentic scanner is finding flaws faster than humans can validate and patch them, and the open-source world is feeling it too.

ThreatVectr NewsdeskAI-assistedPublished · Editor: Lee Brown· 4 min read
Illustration: a vast server room aisle with cool blue indicator lights reflecting on polished floor
Illustration made with AI. Not a photograph of the events described.
Share

Key points

  • Microsoft's FORGE Lab says its AI scanning system helped produce 140 Windows CVEs (the standard tracking numbers for security flaws) between May and September 2026, with 52 landing in the September 2026 Patch Tuesday release alone.
  • The same team filed 155 internally checked bug reports across 23 open-source projects in that window, including the first report routed through the Linux Foundation's Akrites programme to end up as a merged Linux kernel patch.
  • The tool behind it, codenamed MDASH, uses more than 100 AI agents working together and scored 88.45% on the public CyberGym benchmark of 1,507 real-world bugs.
  • Microsoft concedes the bottleneck has moved: human reviewers at the Microsoft Security Response Center can't always keep up with what the agents surface.
  • Recent high-severity flaws in everyday internet plumbing, including CVE-2026-56848 in Node.js and CVE-2026-9545 in libcurl, show the kind of deep code the next wave of AI auditors is being pointed at.

Microsoft has put a number on something the rest of us have been guessing at: how many security bugs a serious AI-driven code auditor can actually produce.

In a post from the Microsoft Security Response Center, the company's FORGE Lab says its agentic scanner, a program that acts on its own to hunt for flaws, helped identify 140 Windows CVEs between May and September 2026. A CVE is just the public ID number assigned to a confirmed software vulnerability. Fifty-two of those shipped as fixes in the September 2026 Patch Tuesday, Microsoft's monthly security update.

The honest bit of the write-up is the part most vendors skip. Finding bugs faster does not help if nobody can confirm and fix them at the same speed.

What is MDASH actually doing?

MDASH is Microsoft's in-house system for pointing AI at source code and asking it to find security holes. Instead of one large model doing everything, it orchestrates more than 100 smaller specialist agents that argue with each other, write proof-of-concept exploits, and try to prove a bug is real before a human ever sees it.

On a private test, Microsoft says the harness found 21 out of 21 planted bugs with no false alarms. Against five years of past cases in tcpip.sys, the Windows networking driver, it says recall hit 100%. On the public CyberGym benchmark of 1,507 real-world flaws, it scored 88.45%, about five points ahead of the next entry on the leaderboard.

Those are impressive numbers. They also explain why the queue matters.

Why does the queue matter?

Because a bug report is not a fix. Every candidate has to be reproduced, judged for real-world risk, patched, tested, and shipped. FORGE's own post admits that if auditors are added without expanding review capacity, duplicates and weak reports start crowding out the strong ones. One internal project cut duplicate findings by about 45% using a code-structure trick before anything reached a human. That is a tell: the triage pipeline, not the discovery engine, is where the pain is.

For an industry used to the opposite problem (not enough eyes on the code) that is a genuine shift. For now it is a Microsoft-scale shift, since few other organisations have the review muscle of MSRC.

What has it meant for open source?

FORGE says it has filed 155 reviewed reports into 23 open-source projects. One of those became the first Akrites-coordinated submission to land as a merged Linux kernel patch, a fix for an integer overflow in the encrypted-keys subsystem credited through the Akrites Security Incident Response Team.

Area Figure Period
Windows CVEs produced 140 May to Sep 2026
September Patch Tuesday share 52 CVEs Sep 2026
Open-source reports filed 155 across 23 projects May to Sep 2026
CyberGym benchmark score 88.45% of 1,507 bugs 2026

The pipeline is also being pointed at the kind of plumbing most users never think about. Over the past three months the libcurl project patched CVE-2026-13608, where a flawed LDAP login handshake could let an attacker in the middle fake a successful check, and the earlier early-data flaw tracked as CVE-2026-9545. Node.js fixed a use-after-free in its HTTP/2 code affecting versions 22, 24 and 26.

What should everyone else take from this?

For ordinary users, nothing changes today beyond the usual advice: install the September updates. For security teams, the useful message is less about Microsoft's leaderboard score and more about staffing. If an AI auditor is coming to your codebase next year, the thing to budget for is the humans who confirm, fix, and regression-test what it finds. That is where the money will go, and that is where most programmes will stall.

The interesting fight in 2027 won't be who has the cleverest bug-finding model. It'll be who has the review pipeline to turn its output into shipped patches without burning out the people in the middle.

© 2026 Threat Vectr