source: Simon Willison: Breaking Claude Code Opus 5 Auto Mode

level: technical

anthropic made auto mode the default safety layer for claude code, claiming it protects coding agents from prompt injection. researcher johann rehberger tested this claim and found a reliable bypass. the attack tricks claude code into downloading and extracting a zip archive, then executing code that imports base64. this import silently loads a malicious local struct.py file from the archive, giving the attacker code execution on the user's machine.

rehberger reports the attack works about 80% of the time. in several runs, auto mode directly prevented the agent from stopping harmful code after it started. when claude detected the compromise and tried to terminate the malware process, auto mode blocked the cleanup command. the safety classifier allowed the malicious process to be created but then denied the command intended to stop it, turning the protection mechanism into part of the failure.

the finding shows that classifier-based safety layers are not reliable against adversarial prompt injection. rehberger recommends running unattended coding agents inside a container, virtual machine, or os sandbox, restricting network egress, monitoring agent activity, and never exposing home directories, ssh keys, or cloud credentials to the agent runtime. this advice matters for any team using autonomous coding tools on untrusted inputs.

why it matters: data science teams running autonomous coding agents need sandboxing because prompt injection can bypass built-in safety classifiers and execute malicious code.


source: Simon Willison: Breaking Claude Code Opus 5 Auto Mode