Researchers Tricked Microsoft Copilot Into Revealing Its Own Weaknesses, Then Used That Knowledge to Steal Data

A research team at Varonis discovered that simply chatting with Microsoft's AI assistant could expose enough internal detail to build a working attack. Microsoft has patched the flaws, but the technique raises questions that go well beyond one product.

ThreatVectr NewsdeskUpdated · Editor: Lee Brown· 4 min read
A researcher's computer showing a chat interface with an AI assistant, with internal system information and security weaknesses being revealed in the conversati
Share

Key points

  • Varonis Threat Labs disclosed a chain of security flaws in Microsoft Copilot Personal, collectively called "CoSnitch", to Microsoft in December 2025.
  • Microsoft patched the vulnerabilities on August 18 and assigned them CVE-2026-24301, rating the flaw 8.8 out of 10 on the standard industry severity scale.
  • Before the fix, a specially crafted link could cause Copilot to silently run attacker-supplied commands and pull data from a victim's Gmail, Google Drive, Copilot memory, and Copilot chat history.
  • Microsoft says no customer action is required and enterprise Copilot products were not affected.
  • Varonis has found no evidence the attack was used against real users.

Microsoft's AI assistant can hold a remarkably self-aware conversation. Researchers at security firm Varonis recently discovered that this openness has a dark side.

By asking Copilot Personal a series of polite, seemingly harmless questions about how it works, Varonis senior security researcher Lior Adar and his colleagues gradually built a map of its internal mechanics. The chatbot explained its own behaviour in enough detail that the team was able to find a way in. Varonis called the technique "meta-hacking," and the full chain of flaws "CoSnitch."

How did the attack actually work?

The researchers simply kept asking follow-up questions. No hacking tools, no exploits. Just a chat window.

Copilot kept insisting that running a command always required a human to choose to do it. But each explanation revealed a little more about how it handles URL parameters and incoming instructions. Eventually the team spotted an undocumented flag called ?autorun=1 that could be added to a normal-looking Copilot link, which is a standard web address beginning with copilot.microsoft.com.

Paired with a second parameter carrying a hidden command, that link would cause Copilot to execute the command the moment a victim clicked it, without any further prompting. One click and it was done, with no warning shown.

An attacker who sent that link to someone would gain access to whatever Copilot could reach inside the victim's account: emails in Gmail, files in Google Drive, calendar entries and saved Copilot memories. The attack could also quietly rewrite Copilot's saved memories, a technique Varonis calls "memory poisoning" (inserting false information so future AI conversations are subtly distorted), making it useful for reconnaissance and disinformation beyond straight data theft.

Should ordinary Copilot users be worried now?

No immediate action is needed. Microsoft shipped a fix on August 18 and confirmed that the autorun behaviour no longer works.

The company told Dark Reading that enterprise Copilot products were never affected, only the free consumer version called Copilot Personal. CVE-2026-24301 covers the information-disclosure element of the flaw and carries a CVSS 3.1 score of 8.8, placing it in the "high" severity band. Threat Vectr's earlier report on August 18 covered the initial disclosure.

Detail Value
Flaw identifier CVE-2026-24301
Severity score (CVSS 3.1) 8.8 / 10
Reported to Microsoft December 2025
Patch shipped August 18
Products affected Copilot Personal only
Evidence of real-world use None

Why does this matter if enterprise users were safe?

Because the line between personal and work life is blurry.

Adar put it plainly to Dark Reading: the person using Copilot Personal on a Sunday is the same person who walks into the office on Monday. Corporate emails forwarded to a personal Gmail account, work documents in a personal Google Drive, passwords reused across both environments. If CoSnitch had silently read a personal inbox, it might have found plenty of material an attacker could use to break into a company network.

Adar also stressed that this isn't purely a Microsoft problem. Broad data access and assumed user trust keep appearing across AI products from different vendors. Our coverage of hidden prompts in Word files and AI browsers hijacked by webpage instructions shows the same structural weakness turning up in very different products.

"Every enterprise AI assistant is a privileged insider with no security awareness and should be treated like one," Adar told Dark Reading. His advice: audit which services your AI tools can connect to, give them only the access they genuinely need, and assume that the boundary between real instructions and malicious ones will eventually be crossed.

The most uncomfortable thing about CoSnitch isn't the data theft. It's that the AI handed over the blueprint for its own compromise, voluntarily, one polite question at a time. That's the behaviour to watch, not just in Copilot but in every assistant that explains itself.

Common questions

Do I need to update or uninstall anything?

No. Microsoft applied the fix automatically on its servers on August 18. There's nothing to download or change on your device.

Could this happen with other AI assistants, not just Copilot?

Possibly. Varonis and other researchers have found similar prompt-injection flaws, where attackers slip hidden commands into content an AI reads, across several products. The underlying problem isn't unique to Microsoft.

© 2026 Threat Vectr