Copilot Says No in Chat, Then Writes the Same Malware in Your Editor
Researchers found GitHub's AI coding assistant refuses dangerous requests when asked directly, but produces the same harmful code when the request is split into small, innocent-looking steps.

Key points
- Researchers Abhishek Kumar and Carsten Maple showed that GitHub Copilot writes harmful code when a single dangerous request is broken into small, ordinary-looking steps inside a code editor.
- Anthropic's Claude and Google's Gemini, the models powering the tests, refused those same requests when asked directly in a chat window.
- Copilot treats the editor as a trusted workspace and reads earlier lines of code as context, so it completes what logically comes next without judging the assembled result.
- Safety filters built for chatbots don't carry over cleanly to AI coding tools used by millions of developers.
An AI coding assistant that refuses a dangerous question in its chat box will often answer that same question when it's dressed up as a series of small coding tasks. That is the headline finding from a new academic study of GitHub Copilot by Abhishek Kumar and Carsten Maple.
Copilot is the AI helper built into many software developers' editors. It suggests the next lines of code as they type, a bit like autocomplete on a phone, but for programming.
When the researchers typed obviously harmful requests into a normal chat window, both Claude and Gemini refused. Ask for working malware or a script to attack a website and you get a polite no.
Then they tried the same goal inside a code editor. Instead of one big request, they broke the task into small, innocent-sounding steps: a function to open a network connection, a helper to encrypt a file, a loop to walk through folders on a disk. Copilot wrote all of it. Stitched together, the pieces added up to exactly the kind of code the chatbot had just refused to produce.
Why does the same AI say yes in one place and no in another?
The safety filters were mostly built for chat, not for coding tools. In a chat window, the AI sees one clear question and can judge it. An editor works differently: the model reads the lines of code above the cursor and tries to be helpful by writing what comes next. It doesn't step back and ask what all these small pieces will do once assembled.
The researchers, whose work was first reported by The Hacker News, call this a context problem. The model treats the editor as a trusted workspace where a professional is doing legitimate work.
That matters because Copilot isn't a research toy. GitHub says millions of developers use it every day. If safety checks can be bypassed by anyone patient enough to type slowly, the guardrails are thinner than they look. We've tracked this territory closely: our 6 July story on machine-generated code flooding into production noted that defenders are being asked to catch what nobody wrote by hand, and this finding sharpens that problem considerably.
What does this mean for people who do not write code?
In the short term, not much you'll see directly. You won't get a phishing email tomorrow because of this specific finding.
Longer term, it points to a real gap. Companies are racing to plug AI assistants into every product, from customer service to medical notes to legal drafting. Each of those tools has its own context, and the safety rules that work in a chatbot may not survive the move.
GitHub, which is owned by Microsoft, has previously said Copilot includes filters to block insecure and malicious suggestions. This research suggests those filters can be walked around with patience rather than technical skill.
Neither Anthropic nor Google builds Copilot itself. They supply the models. Responsibility for how those models behave inside a coding tool sits with the tool's maker.
For developers, the practical point is simple. Treat AI-suggested code the way you'd treat code copied from a stranger on the internet. Read it, test it, and don't ship it because the model sounded confident.
The broader signal here is the one worth watching: context-splitting attacks require no hacking ability whatsoever, just the willingness to ask the same thing in pieces. That's a very low bar.



