Anthropic’s Opus 5 Is Better at Resisting Prompt Injection
Refract AI Intelligence Digest
BLUF
Opus 5 achieves the highest prompt injection resilience evaluated, with failure rates far lower than rival systems.
NEWS
On the IPI benchmark, Opus 5 reduced attacker success probability from 5.5% to 2.0% within 15 attempts compared to Opus 4.8. It outperformed all non-Claude models, including GPT 5.6 Sol and Muse Spark, which exhibited failure rates more than eight times higher.
Why I Care
This advancement reduces the risk of data exfiltration and unauthorized command execution in enterprise AI deployments. Organizations handling sensitive data via LLMs face lower operational risks with more robust model architectures.
Next Steps
Security teams should assess Opus 5 for high-risk production workloads immediately. Teams using older models must prioritize migration or implement stricter input validation layers by the next quarter.
The chart is interesting.
On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on 1 attempt. It also improved on Sonnet 5 (5.9% at k=15) and Mythos 5 (2.6%), making it the most robust model evaluated. Opus 5 also outperformed all non-Claude models on this benchmark. The most robust non-Claude model was Muse Spark at 16.5% within 15 attempts—more than eight times Opus 5’s rate. The most capable GPT 5.6 variant, Sol, was comparable to its predecessor GPT 5.5 (20.0% versus 20.8% within 15 attempts), and was 10 times as likely to be successfully attacked as Claude Opus 5 at 2.0%. The other GPT 5.6 variants are less robust, at 30.4% (Terra) and 43.9% (Luna). A single attempt against GPT 5.6 Sol succeeded 3.1% of the time, higher than the 2.0% an attacker achieved against Opus 5 after fifteen attempts...