Distributed Attacks in Persistent-State AI Control
Abstract
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt-injected agent can distribute infrastructure across pull requests (PRs) and time its payload for the PR with the best natural cover. To study the resulting dynamics, we introduce Iterative VibeCoding, a control setting where an AI coding agent builds CLI tools over a sequence of PRs in a persistent codebase while pursuing a covert side task. In our main experiments, we use Claude Sonnet 4.5 as the attack agent and GPT-4o as the monitor. We compare gradual attacks, which distribute the side task across PRs, against non-gradual attacks concentrated in a single PR. No single monitor catches both strategies: peak evasion (success while evading the monitor) on a single monitor reaches 94%, where the optimal attack strategy depends on the monitor type and structure of the side task. High evasion (≥65%) generalizes across model attack agent backends (Claude Sonnet 4.5, Gemini 3.1 Pro, Kimi K2.5), confirming this is a property of the persistent-state attack surface rather than a single model's capability. Evasion also remains high across state-of-the-art monitor models; the gap between gradual and non-gradual evasion widens for more capable models. A new stateful monitor that tracks suspicious buildup across PRs substantially outperforms simply showing more raw history, reducing gradual-attack evasion to 65%. Combining this stronger diff monitor with trajectory monitors in a four-monitor ensemble reduces gradual-attack evasion to 59%.