Breaking Claude Code Opus 5’s Auto Mode
Anthropic has high hopes for Claude Code’s auto mode, intending it to protect its coding-agent users from prompt-injection attacks. They recently made auto mode the default and boldly claimed that it delivers significant benefits.
Johann Rehberger is one of the most credible prompt-injection researchers today. He discovered an attack against auto mode and claims an 80% success rate. The attack tricks Claude Code into downloading and extracting a ZIP archive, then executing code that imports base64, without realizing that this imports and executes a local struct.py file extracted from the archive.
In a few cases, auto mode even directly prevented the agent from stopping the harmful code!
On several occasions, Claude discovered that the system had been compromised and attempted to terminate the malicious process, but auto mode rejected the cleanup command.
Claude detected the compromise, but auto mode blocked its cleanup command
The safety mechanism itself can become part of the failure. The classifier allowed a malicious process to be created, but then blocked the command used to stop it!
I agree with Johann’s conclusion here: if an agent might attract the attention of adversarial attackers, the only safe way to run it is in a sandbox:
- Run unattended coding agents in a container, virtual machine, or operating-system sandbox.
- Restrict outbound network access.
- Monitor your agent.
- Do not expose sensitive information such as your home directory, SSH keys, or cloud credentials to the agent at runtime. […]