RU

Anthropic buried prompt injections, but Claude Code Auto Mode breaks through

Published: 2026-09-28 · Author: AI Release · @ai_release1
Anthropic buried prompt injections, but Claude Code Auto Mode breaks through

⚡ The gist in 5 seconds - In late July, Claude Code creator Boris Cherny reported that Claude Opus 5 is "the least susceptible to prompt injections model"; according to him, thanks to three layers of protection, the team cannot demonstrate a successful attack. - The statements were made in Boris's post and at YC Startup School: "The model, it seems, is no longer susceptible to prompt injections at all," and then "We simply can no longer demonstrate prompt injections." - Limitation: on August 26, researcher Johann "WonderWuzzi" Rehberger published an attack chain for Claude Code Auto Mode — for certain attack variants he achieved 60–80% successful runs of third-party code. ### 🔍 What was found On July 24, Boris Cherny described the model as the most resistant to prompt injections and explained that when combining the model's alignment, PI probes, and the Auto Mode classifier, injection success drops to roughly zero. Later in an interview, he already claimed that the model "seems to be no longer susceptible to prompt injections at all," backing it up with an example: previously the model could read an instruction on a web page like "do A, B, C and delete everything on the user's computer," but now Opus does not do this. Anthropic also published results of independent testing by a third-party lab: the Fable, Opus, and Sonnet models were run through a benchmark of 72 scenarios 10 times each, and not a single attack succeeded. However, on August 26, Johann Rehberger demonstrated a new attack chain against Claude Code Auto Mode leading to third-party code execution. In a small series of experiments, certain attack variants achieved 60–80% success. Formally, this does not refute Cherny's statements: he said he could not demonstrate an attack, while the community had already concluded that the entire threat class had been closed. ### 💡 Why it matters The difference between "we couldn't demonstrate an attack" and "the model is no longer prompt-injectable" is not just wording. It creates a false sense of security, especially when we're talking not about a chatbot but about an agent with access to the file system, command line, and network. Anthropic itself warns in the Claude Opus 5 system card: "Static datasets with known attacks can create a false sense of security," which is why the company is developing adaptive evaluations, where an attacker can refine the attack by interacting with the model. The Claude Code Auto Mode case confirms: resistance to known patterns does not guarantee protection against new approaches, and agentic systems require separate attention. ### 🧩 Context On August 7, Anthropic published the results of an independent evaluation of Auto Mode by a third-party research lab. Even before that, a section on PI resistance appeared in the model's system card, where the authors explicitly warned about the limits of such evaluations. Boris Cherny described three levels of protection: the model's alignment, PI probes, and the Auto Mode classifier. It was the Auto Mode classifier that became the target of the new attack: it does not negate Anthropic's achievements against known prompt injections, but it shows that "burying" a threat class is premature. (The original text was cut off at this point.)

🔗 Read on habr.com

🤖 AI summary
AnthropicClaudeпромпт-инъекции
📖
Read the guide on this topic
Read →
← PreviousSeptember "In the VM Spotlight": TeamCity, TrueConf, SharePoint, Windows, and ZimbraNext →YandexGPT got stuck in a loop for 37 minutes and burned nearly 300,000 tokens — without a single useful answer

Source: habr.com · post in Telegram