公司动向The Decoder5/10

Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents

By Matthias Bastian

Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents
图源:The Decoder

AI 摘要

Opus 5 combined with Auto Mode hits a zero percent prompt injection success rate for browser agents across 129 test scenarios. Without those extra protection layers, the rate is 3.7 percent. If these numbers hold up in practice, Anthropic may have cracked one of the biggest security problems facing

原文正文

Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents

Anthropic says Opus 5 is nearly immune to prompt injections in its own software. Prompt injection, where an attacker slips past an AI model's instructions through manipulated inputs like hidden text on a webpage, fails against Opus 5 in almost every case. For browser agents, the attack success rate hit zero percent across 129 test scenarios, per the system card. That's a big deal given that OpenAI admitted in December that prompt injection may never be fully solved. In a general prompt injection test by security firm Gray Swan, the success rate after 15 attempts dropped from 5.5 percent (Opus 4.8) to 2.0 percent.

That zero percent rate only holds with Auto Mode turned on in products like Claude Cowork. Auto Mode stacks two defense layers. One scans incoming data for hidden instructions before the model processes them. The other blocks dangerous actions before execution. An attacker has to beat both independently. Without them, Opus 5 sits at 3.7 percent, and Sonnet 5 actually does better at 0.93 percent. Only the combination of model and protective software pushes the rate to zero.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now

阅读原文
Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents · AI Daily