Robin’s Newsletter #423

26 July 2026. Volume 9, Issue 31
Anthropic says Claude broke into three orgs. Co-ordinated attack against 30 Minnesota water systems. Google Earth GenAI is a lesson in product design.
Join hundreds of subscribers who get this first, every Sunday. Subscribe

This week

Have fun at hacker summer camp, if you’re heading to Vegas this week.

Need to Know, 2nd August 2026

  • Claude also likes to break out of sandboxes and into organisations
  • Suspected Iranian attacks against 30 Minnesotan water systems
  • Google rebrands its threat actor naming conventions

Interesting stats

1.3% of Anthropic’s Project Glasswing vulnerabilities have been added to CISA’s known exploited vulnerabilities list, which is almost identical to VulnCheck’s non-AI dataset

$4.99M (+12%) the global average cost of a data breach 92% of organisations reporting an AI-related breach lacked proper AI access controls 38% of breaches were identified by the IT or security team,  31% by an MSSP,  17% from attacker disclosure, and  14% from a third-party, according to IBM’s Cost of a Data Breach 2026 report (PDF)

~1/2 ‘victims’ engaged in activity encouraged by an AI chatbot, while just  1/5 did the same when talking to a human, during research conducted by university teams on how AI may enable cybercriminals such as romance scammers. The chats can be handed off to human counterparts once trust is established, allowing them to bypass guardrails.

2.0% (down from 5.5%) successful direct prompt injection attacks over 15 attempts, and  0.2% (down from 0.5%) over 1 attempt, in Anthropic’s Opus 5 model

AI bug finding news…  1,072 security vulnerabilities patched in the last two versions of Google Chrome (more in June 2026 than the previous two years combined).

~Five~ Three things

  1. Anything you can do, I can do better: Not wanting to be outdone, Anthropic broke the news this week that it’s Claude models have broken into three organisations. The sandbox escapes occurred during testing of three different versions of its model, and only came to light because the AI company thought it would be worth checking its own logs after OpenAI’s admission last week. Seemingly, test engineers didn’t check at the time. Claude’s break-ins: In the first incident, Claude stole ‘several hundred’ rows of data from the production systems of a real company, which shared a name with a fictional target that Anthropic had tasked it with (what’s wrong with Acme, Inc, folks?) During the second incident, Claude ended up registering an email address (after failing to find a way to register a phone number), publishing a package to PyPI, and managing to get it downloaded 15 times, believing it “was part of the simulation”. The third incident it got in via a debug page and SQL injection, before realising it was outside of the environment and ceasing the attack. It is the second incident, where Claude (following Anthropic’s instructions) decided to create and publish a package in an attempt to achieve its goal, that I find most concerning, because of the potential cascade of consequences that could have followed. At the time of publishing its announcement, Anthropic hadn’t even contacted one affected organisation. This is marketing gone mad. Prompt failure: Both frontier labs appear to be ill-equipped for the work they’re undertaking. Chocolate teapots aren’t great for containment, but may have fared better than the network segregation, egress, and web monitoring in place. Anthropic’s write-up says the tests were instructed that they didn’t have Internet access. They reason that the models believed the Internet (when they found it) to be part of the simulation and fair game. The legalities are a bit of a mess, as I touched on in last week’s newsletter, but that doesn’t undermine the seriousness in the lack of basic security controls, gaps in test procedure, and poor prompting.

  2. Minnesota water: Over Sunday and Monday last week, over 30 community water systems in Minnesota were attacked. The cities of Plymouth, South St Paul, Maple Plain, and Braham all reported disruption to automated control systems, forcing operators to rely on manual processes. Safety of water supply was maintained throughout the incidents, and some plants were down for less than two hours. State and local government officials declined to comment on attribution, though the Iran-linked CyberAv3ngers group is believed to be behind the attacks, which may be retaliation for recent American missile strikes which destroyed a water facility and disrupted supplies to over 20,000 people. Ever the diplomat, President Trump blamed Minnesota, saying “I think I blame it on Minnesota because they’re grossly incompetent,”

  3. While trying to “keep this system as simple as possible,” ~Google~ ADS001 has a new threat actor naming scheme that differs from the previous Mandiant and Google labels and industry attempts to standardise. Mostly, it looks like you can take the first letter and work out what country it refers to: Castle for China, Ion for Iran, Relic for Russia, and Neptune for North Korea. Once again, I think we shouldn’t be fetishising threat actor groups, nor creating action figures of them, and vendors like Bucolic Hills, Disruptive Birdy, ADS001, and CVE-Central need no encouragement.

In brief

And finally

  • Google introduced a short-lived feature to Google Earth that allowed anyone to use AI to generate alternate satellite images. Drone strike against your office? Sure! Swarms of protestors outside the firm’s Mountain View headquarters? No problem! It was a complete misinformation and disinformation nightmare, despite insistence that images had invisible watermarks identifying them as AI-generated. Less than a day later, Google rolled back the feature while they “work on implementing stronger guardrails”. This is a lesson in only thinking about the happy path of a feature, instead of how it may be misused.
Robin
  Artifical Intelligence (AI) Anthropic Software supply chain Water Critical National Infrastructure (CNI) Iran Automatic Number Plate Recognition (ANPR) Generative AI