Your AI Coding Bots Are Running Unsupervised and Nobody Knows What They Did Last Night

AI agents inside software development teams can write, test, and deploy code on their own, often with no human checking what they did. Most companies have no way to answer a simple question: who authorised that change?

ThreatVectr NewsdeskAI-assistedPublished Updated · Editor: Lee Brown· 4 min read
A dimly lit server room at night, rows of glowing blue and green indicator lights on rack-mounted hardware stretching into the distance, empty operator chairs i
Illustration made with AI. Not a photograph of the events described.
Share

Key points

  • The average organisation runs 22 separate AI agent projects, according to recent industry research, spanning departments from legal to engineering.
  • AI agents in software teams can write code, open pull requests, and push updates to live systems with no human reviewing a single step.
  • Unlike a service account or API token, an AI agent sets its own path to reach a goal, making its behaviour fundamentally harder to predict or audit.
  • Most companies cannot tell regulators or auditors which AI agent introduced a specific piece of code, or whether any human was involved.
  • The compliance rules under the Sarbanes-Oxley Act, a US law requiring companies to prove financial and technical controls, have existed for over 20 years, yet AI agents are making those controls nearly impossible to demonstrate.

This crisis doesn't show up in breach headlines. It shows up in postmortems, audit findings, and 2 a.m. Slack messages asking who approved that deployment.

The problem is AI agents.

Not the helpful kind that suggests a line of code while a developer types. The other kind: fully autonomous programs that receive a goal, figure out the steps themselves, write the code, run the tests, ship the result to production, all while your security team is asleep and your audit log shows nothing useful.

How did these agents end up with so much access?

They got it the same way every other automated tool does: a developer needed something done fast, spun up an account or access key for the agent, and moved on. The access was never revoked. What it touched the following week, nobody checked.

That's the failure mode. It isn't dramatic. Just the ordinary accumulation of ungoverned access, except the thing holding that access now makes decisions independently, at machine speed, across an entire codebase.

A normal service account or API token (a digital pass that lets software talk to other software) does exactly what it's told. An AI agent is handed a target and works out the route itself. That distinction sounds abstract until you realise it means no human defined the specific actions the agent would take, so nobody can easily reconstruct them afterward.

Dark Reading recently covered analysis from identity security researchers who put it plainly: governance models built for human developers assumed human speed and human judgment. Models built for automated tools assumed predictable, bounded behaviour. AI agents break both assumptions at once.

We've been tracking this problem since our 2 July piece on identity governance, and the picture keeps getting worse. The coding-assistant phase, tools like GitHub Copilot or Cursor where a person still makes the final call, is one risk level. Fully autonomous agents are a different category, arriving faster than most governance teams are ready for.

The organisations furthest behind aren't the ones ignoring AI. They're the ones who enthusiastically adopted it, let developers spin up agent projects on personal accounts, and never built the inventory to know what exists.

Auditors and regulators will eventually ask: if an AI agent introduced a vulnerability or made an unauthorised change, can you show exactly what happened, which identity did it, and whether a human approved it? For most organisations today, the honest answer is no.

Should you worry about your CI/CD pipeline specifically?

Yes. The development environment is where AI agents are most embedded, most autonomous, and most ungoverned. Four things help close the gap: know every agent running in your environment, including the unofficial ones; know what each agent can access and what it's actually used; build a baseline for normal behaviour so abnormal behaviour stands out; revoke access when a project ends.

None of that is exotic. It's standard practice for managing human accounts. The frameworks just haven't caught up to apply it to AI agents yet.

The plain judgement here: the code an AI agent produces is the symptom. The ungoverned identity behind it is the cause. Without knowing the cause, you're patching symptoms indefinitely, and one of those symptoms is eventually going to land in front of a regulator.

Operational takeaway: if you can't pull a report today showing every AI agent with write access to your CI/CD pipeline (the automated system that builds and ships software), that report is your first priority next sprint.

© 2026 Threat Vectr