← Back to feed
AI SecurityEmerging1 sourceAug 27, 2026 · 04:00via Embrace The Red (AI agent security)

Breaking Claude Code Opus 5 Auto Mode

Brief

In this post, we explore how a simple website summary request hijacks Claude Code Opus 5 in Auto Mode and achieves code execution with 60-80% attack success rate using a small sample size.

This is interesting because a third-party evaluation commissioned by Anthropic showed a 0.00% prompt injection attack success rate for Opus 5 in Auto Mode.

Auto Mode Is Now the Default in Claude Code

Auto Mode replaces human approval prompts with a safety classifier. Since mid-August it is the default starting mode for Claude Code.

Read more on Embrace The Red (AI agent security)