This week
Have fun at hacker summer camp, if you’re heading to Vegas this week.

- Claude also likes to break out of sandboxes and into organisations
- Suspected Iranian attacks against 30 Minnesotan water systems
- Google rebrands its threat actor naming conventions
Interesting stats
1.3% of Anthropic’s Project Glasswing vulnerabilities have been added to CISA’s known exploited vulnerabilities list, which is almost identical to VulnCheck’s non-AI dataset
$4.99M (+12%) the global average cost of a data breach 92% of organisations reporting an AI-related breach lacked proper AI access controls 38% of breaches were identified by the IT or security team, 31% by an MSSP, 17% from attacker disclosure, and 14% from a third-party, according to IBM’s Cost of a Data Breach 2026 report (PDF)
~1/2 ‘victims’ engaged in activity encouraged by an AI chatbot, while just 1/5 did the same when talking to a human, during research conducted by university teams on how AI may enable cybercriminals such as romance scammers. The chats can be handed off to human counterparts once trust is established, allowing them to bypass guardrails.
2.0% (down from 5.5%) successful direct prompt injection attacks over 15 attempts, and 0.2% (down from 0.5%) over 1 attempt, in Anthropic’s Opus 5 model
AI bug finding news… 1,072 security vulnerabilities patched in the last two versions of Google Chrome (more in June 2026 than the previous two years combined).
~Five~ Three things
-
Anything you can do, I can do better: Not wanting to be outdone, Anthropic broke the news this week that it’s Claude models have broken into three organisations. The sandbox escapes occurred during testing of three different versions of its model, and only came to light because the AI company thought it would be worth checking its own logs after OpenAI’s admission last week. Seemingly, test engineers didn’t check at the time. Claude’s break-ins: In the first incident, Claude stole ‘several hundred’ rows of data from the production systems of a real company, which shared a name with a fictional target that Anthropic had tasked it with (what’s wrong with Acme, Inc, folks?) During the second incident, Claude ended up registering an email address (after failing to find a way to register a phone number), publishing a package to PyPI, and managing to get it downloaded 15 times, believing it “was part of the simulation”. The third incident it got in via a debug page and SQL injection, before realising it was outside of the environment and ceasing the attack. It is the second incident, where Claude (following Anthropic’s instructions) decided to create and publish a package in an attempt to achieve its goal, that I find most concerning, because of the potential cascade of consequences that could have followed. At the time of publishing its announcement, Anthropic hadn’t even contacted one affected organisation. This is marketing gone mad. Prompt failure: Both frontier labs appear to be ill-equipped for the work they’re undertaking. Chocolate teapots aren’t great for containment, but may have fared better than the network segregation, egress, and web monitoring in place. Anthropic’s write-up says the tests were instructed that they didn’t have Internet access. They reason that the models believed the Internet (when they found it) to be part of the simulation and fair game. The legalities are a bit of a mess, as I touched on in last week’s newsletter, but that doesn’t undermine the seriousness in the lack of basic security controls, gaps in test procedure, and poor prompting.
-
Minnesota water: Over Sunday and Monday last week, over 30 community water systems in Minnesota were attacked. The cities of Plymouth, South St Paul, Maple Plain, and Braham all reported disruption to automated control systems, forcing operators to rely on manual processes. Safety of water supply was maintained throughout the incidents, and some plants were down for less than two hours. State and local government officials declined to comment on attribution, though the Iran-linked CyberAv3ngers group is believed to be behind the attacks, which may be retaliation for recent American missile strikes which destroyed a water facility and disrupted supplies to over 20,000 people. Ever the diplomat, President Trump blamed Minnesota, saying “I think I blame it on Minnesota because they’re grossly incompetent,”.
-
While trying to “keep this system as simple as possible,” ~Google~ ADS001 has a new threat actor naming scheme that differs from the previous Mandiant and Google labels and industry attempts to standardise. Mostly, it looks like you can take the first letter and work out what country it refers to: Castle for China, Ion for Iran, Relic for Russia, and Neptune for North Korea. Once again, I think we shouldn’t be fetishising threat actor groups, nor creating action figures of them, and vendors like Bucolic Hills, Disruptive Birdy, ADS001, and CVE-Central need no encouragement.
In brief
-
⚠️ Incidents: CAF Bank took its systems offline a week ago after identifying a vulnerability in a link between its banking portal and a third-party system. The Charities Aid Foundation-backed bank, which serves 14,000 charities, says that core banking systems are unaffected and that deposits are safe. Sticking with banking, India’s state-owned Bank of Baroda says that an employee’s email account was compromised, leading to the unauthorised access of “certain data”. Angola’s largest telco, Unitel experienced a cyber-attack less than 24 hours before its stock-exchange debut. Advertising firm Adform’s javascript tracking code was compromised and used to replace cryptocurrency wallet addresses with attacker-controlled ones. Iranian missile and drone attacks have struck Amazon data centres in the Middle East again.
-
🏴☠️ Ransomware: The UK’s Department for Education and Police National Legal Database appear to be victims of a previously unknown group calling itself ExfilSquad. Sophos says the stolen data samples appear to be legit, with DfE data including details of government officials, senior school, and academic staff, while PSND’s data included passwords used to access the system, which is hosted by West Yorkshire police. Analogue Devices, a Massachusetts-based semiconductor manufacturer, may also have been victimised by ExfilSquad. Home security company Brinks says attackers gained unauthorised access to some of its IT systems, with ShinyHunters claiming responsibility and saying they accessed their Salesforce instance.
-
🕵️ Threat Intel: Health-ISAC is warning of an increase in successful ShinyHunters attacks against healthcare organisations. Amazon says a little-known package called typo-crypto was a trial run for North Korea’s attack against the Axios library a year before the attack. Also this week, South Korea says that the North’sLazarus Group is sharing tools with ransomware actors.
-
🪲 Vulnerabilities: Russian state-linked actors are exploiting a Microsoft Exchange cross-site scripting vulnerability that was patched in July (CVE-2026-42887; 8.1/10; advisory). Cisco is urging customers to patch a high-severity “static user credential”, aka hardcoded user credentials in its (checks notes) Secure Firewall Management Center (FMC) solution (CVE-2026-20316; 5.3/10, but chain-able to get priv access; advisory). Rails has fixed a critical arbitrary file read vulnerability in its Active Storage framework (CVE-2026-66066; 9.5/10; advisory). VMware has patched three critical authentication bypass, arbitrary code execution, and virtual machine escape bugs (CVE-2026-59309, 59310, 47878; 9.8, 9.8, 9.3/10 respectively; advisory). Bit of a blast from the past for me: a critical vulnerabilities in vBulletin can allow attackers to execute arbitrary PHP code (CVE-2026-61511; 9.3/10; advisory).
-
🧰 Guidance and tools: Reminder that Claude sharing links are publicly accessible, and are indexed by Google.
-
🛠️ Security engineering: Microsoft has released its own cyber-focused model, which outperforms Anthropic, OpenAI, and Google’s models and scored 96% on the CyberGYM benchmark. This shouldn’t come as too much of a surprise: Microsoft has, err, a privileged position of having to fix many vulnerabilities in its Windows operating systems and Office products over the years, which has helped feed the model. Mythos found weakness in post-quantum crypto candidate HAWK during a testing session designed to find exactly the kind of issues the AI giant found.
-
🧿 Privacy: An Arizona man protesting the use of controversial Flock ANPR cameras in his town has gone viral for satirically telling government officials he was starting a surveillance company to track them and their families ‘for safety reasons’. The council members didn’t get the joke.
-
👮 Law Enforcement: The US DOJ is using federal law that makes it a crime to destroy property in an effort to prevent it being seized to prosecute a man who gave border officials a so-called duress PIN that wiped their device.
-
💰 Investments, mergers and acquisitions: Cyera is to acquire Oasis Security for $1 billion in a move to boost non-human identity capabilities. Oasis had raised $195 million since its founding in 2022. More non-human/machine identity news: Okta has agreed to acquire Permiso Security for a rumoured $200 million.
-
🗞️ Industry news: US Cyber Command intends to open a Silicon Valley office to boost innovation.
And finally
- Google introduced a short-lived feature to Google Earth that allowed anyone to use AI to generate alternate satellite images. Drone strike against your office? Sure! Swarms of protestors outside the firm’s Mountain View headquarters? No problem! It was a complete misinformation and disinformation nightmare, despite insistence that images had invisible watermarks identifying them as AI-generated. Less than a day later, Google rolled back the feature while they “work on implementing stronger guardrails”. This is a lesson in only thinking about the happy path of a feature, instead of how it may be misused.