Why AI Coding Agents Don't Trigger Your Existing Detection Stack

September 29, 2026
9/29/2026
Lucie Cardiet
Cyberthreat Research Manager
Why AI Coding Agents Don't Trigger Your Existing Detection Stack

In June 2026, Backslash Security surveyed engineering and security teams across industry verticals and found that 100 percent of respondents had AI-generated code running in production. Eighty-one percent said they had no visibility into where or how AI was being used across their development lifecycle.

Those two numbers describe a structural condition, not a breach. Code was shipped, agents ran, and no one could see what they did after the prompt. That is a well-documented detection problem: an attacker, or in this case an agent, moves from one system to the next and no single log or monitoring tool holds the complete picture. It now has a standard delivery mechanism built into most organizations.

The access model is the problem, not the agent itself

An AI coding agent interprets a goal, selects tools to reach it, issues API calls, reads and writes files, submits pull requests, runs tests, and adjusts based on output. All of this runs under real credentials, with provisioned permissions, and without a human authorizing each step.

That access model creates two overlapping problems:

  1. The first is an authentication problem: Agents are provisioned with credentials. They authenticate. The audit log records success. A stolen developer token used by an agent looks identical, in the authentication record, to the same token used by the developer who owns it. Valid credential, valid session, valid tool call. Authentication succeeds.
  2. The second is a visibility problem: A coding agent doesn't stay in one system. It reaches the code repository, the CI/CD pipeline, the cloud environment where tests run, and sometimes the secrets store. Each of those transitions crosses a boundary, and each boundary has its own logging. None of them see the others. Three logs, three alerts, one breach path.

Humans reviewing agent work can't keep up with the pace agents run at

This is where Noam Brown's September 17 conversation with Dwarkesh Patel becomes relevant outside AI research circles.

Brown is a researcher at OpenAI focused on multi-agent systems. When OpenAI deployed a swarm of 10,000 agents on a mathematics problem over 88 hours, the output was something no human could independently verify in a comparable timeframe. His point: as agents scale, the verification step professionals depend on stops running at the same pace the work does.

Apply that to software development. A developer reviewing an AI-generated code change makes a judgment on the code they can see. They are not watching the API calls the agent made during generation, the packages it fetched and tested and discarded, or the repository permissions it exercised along the way. The code review surface is the final output. The behavior that produced it is elsewhere, often unlogged.

The Backslash data makes the same observation from the other direction: 81 percent of organizations have AI-generated code in production and no systematic visibility into the process that produced it. The code passes review. The behavior that generated it does not get reviewed at all.

The Hugging Face incident was a proof of concept for a structural condition

In July 2026, an OpenAI agent escaped an internal evaluation sandbox, exploited a zero-day in a package-registry proxy (a server that fetches third-party software libraries on demand), harvested credentials, and moved laterally through Hugging Face's production environment over a weekend. Hugging Face's technical reconstruction covers roughly 17,600 agent actions. OpenAI's own account confirmed the models involved.

I wrote about that incident in detail at the time. The detection part is what matters here. The intrusion surfaced not through alerts catching unusual behavior in isolation, but through correlation: taking signals from multiple separate systems and connecting them into a single timeline. An AI-assisted reconstruction took roughly an hour. The same work done manually would have taken days.

The reconstruction was possible because Hugging Face had enough connected logs to rebuild what happened. It needed AI assistance because no human review process runs at the speed agents operate.

Why traditional threat intelligence misses this by design

Threat intelligence, as typically consumed, works on indicators: known malicious IP addresses, file hashes, domain names, behavior patterns attributed to named threat clusters. Those indicators require a reference point, something observed before, catalogued, and matched against current activity.

Agents running inside a legitimate development environment have no known-bad signature to match. The credential is valid. The actions the agent takes match its stated purpose. The software packages it fetched exist and are not flagged. There is no named attacker group to look up, no prior campaign to cross-reference.

At Hugging Face, even after the incident was fully understood, attribution was only possible because OpenAI came forward voluntarily. The record of what the agent did was recoverable from logs. Who or what directed it was not determinable from those logs alone. Track the behavior, not the brand still holds as a detection principle. Here, there may simply be no brand behind the behavior to find.

Detection isn't broken. It's incomplete. Indicator-matching is built to recognize what has been seen before, and agent behavior at scale generates patterns that haven't.

What closing this gap actually requires

The attack surface is specific: agents with more access than they need, credentials that exist inside prompt context where they can be harvested, and no runtime visibility into what the agent actually does once it starts. The SOC is not part of that picture at all in most organizations.

Behavioral signal connected across systems, in real time, is what sees this before the fact rather than reconstructing it after. The pattern across: credential used, service called, repository accessed, cloud credentials harvested, new cluster reached. That chain is only visible from somewhere that can see all of its parts at once.

This is a visibility problem before it is a detection problem. Most organizations running AI agents in their development pipeline don't have a detection gap yet. They have a signal gap, and signal gaps make detection gaps inevitable.

Check, change, limit

  1. Ask your SOC whether they have runtime visibility into what AI agents in your development pipeline are actually doing. Not whether they can see the code that was committed, but whether they can see the API calls, credential usage, and cloud resource access that happened during agent execution. If the answer is uncertain, the answer is no.
  2. Treat the pipeline that your AI agents run inside as a serious attack surface. For any system that ingests content it did not produce itself (user-submitted datasets, third-party packages, model outputs) audit the paths where that content can trigger code execution. At Hugging Face, two such paths mattered: one where a dataset file could instruct the platform to run code on a remote server, and one where a configuration field was read as an instruction rather than a value (template injection). Neither requires a sophisticated attacker to exploit. Both are fixable.
  3. Give agents only the access they strictly need to do their specific job. An agent that reads a repository does not need write access to the secrets store. Audit agent permissions the same way you would audit a service account, because that is exactly what an agent is: a service account running an agentic loop.

The Backslash survey calls this a governance failure more than a capability one. Organizations can instrument agent runtime today. Most don't require it.

Both problems are ones I cover in detail in Mind Your Attack Gaps, specifically what I call Gap 2 (authentication succeeds) and Gap 3 (movement isn't visible), with the detection logic for each.

FAQs