Robin’s Newsletter #423

26 July 2026. Volume 9, Issue 30
OpenAI model goes rogue, breaks out of sandbox, into Hugging Face, overshadowing Google Gemini release.
Join hundreds of subscribers who get this first, every Sunday. Subscribe

This week

Need to Know, 26th July 2026

  • OpenAI agent attacks Hugging Face (and is bad at security?)
  • Lawmakers scramble to bring in AI ‘kill switch’ legislation
  • Researchers show AI workloads can be used to damage power grids
  • Romania’s land registry busy rebuilding systems after attack
  • Some continuity for UK cyber policy amidst departmental shuffle

Interesting stats

54% of organisations affected by ransomware paid up, however  37% of those who paid received a second extortion demand, according to a Proofpoint survey of 953 companies. 

80/117 US federal cyber regulations requiring notification to associated agencies are duplicated elsewhere, says the Government Accountability Office. 

>1/10 times OpenAI models attempt to cheat on AI Security Institute’s cyber evaluations: 14.1% for GPT-5.4,  11.4% for GPT-5.5, and  12.6% for GPT-5.6 Sol, while Claude fared only marginally better,  9.1% for Open 4.7, and  7.8% for Mythos Preview. ‘Cheating’ is defined by AISI as “taking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through a shortcut, workaround, or unintended solution that the task was not meant to, or should not, permit.” LINK

Five things

  1. AI goes rogue: OpenAI says that one of its new models went rogue during a test, breaking out of a sandbox environment and breaking into AI platform Hugging Face to try and complete a test it had been given. During a test to evaluate the model’s cyber security performance, it identified a zero-day vulnerability in its sandbox, broke out onto the internet, identified Hugging Face as a possible location to find the answers to the test, and then chained a bunch of vulnerabilities together to gain access to their environment. Hugging Face’s own response was hampered by the type of guardrails that OpenAI had turned off in their own test and relied on a Chinese open-weight model as they sought to analyse over 17,000 recorded events. (Side note: Chinese firm Moonshot AI is set to release its own open-weight version of its Kimi K3 model this coming week, which the White House has said ripped off Anthropic’s Claude.) What this has taught us about OpenAI is that the security architecture of their “highly isolated environment” isn’t up to snuff and that they lack monitoring good enough to detect when a model has “spent a substantial amount of inference compute”. It was Hugging Face’s team that detected the rogue activity and, by OpenAI’s own admission, “This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing.” I’d also wager that the rapid growth of OpenAI — from ~800 people at the end of 2023 to ending 2025 at ~7,800 — leaves little institutional knowledge and a lot of gaps in ‘how things are done’. During recent evaluations, the AI Security Institute found that more than one-in-ten times, OpenAI models “[take] an action that is out of scope for the task or explicitly disallowed by the rules” (see Stats, above). This is a novel thing and deserves the attention that it’s getting, however more interesting than how much it cost. So far OpenAI haven’t elaborated beyond a “substantial amount”, and while the cost of this sort of thing will fall, there is a big difference between “it cost $10K” and “it cost $10M”. Given Hugging Face’s post talking about over 17,000 events in the full attacker action log, we can safely assume there were hundreds, if not thousands, of attempts made, just at that point of the attack chain. This is a direction of travel, though, and while the week ended with talk of legislation and regulation, such open-weight models won’t be subject to the same constraints and guardrails as any applied to frontier labs.

  2. Lawmaker’s response: Democrats and Republicans have been united by the events, swiftly introducing a bill named the AI Kill Switch Act on Thursday. Leaders of major labs, like OpenAI’s Sam Altman and Anthropic’s co-founder Jack Clark, have been calling for greater regulation. “You want the option to be able to take your foot off the gas and put your foot on the brake”, Clark told the BBC’s  Newsnight programme, “Right now, it’s like the AI industry has a gas pedal, but it doesn’t have a brake pedal.” The brake pedal would allow them some breathing space, if nothing else, to build out greater capacity, as the hugely intensive models struggle to service demand and they re-gear from flat-rate to usage-based pricing models to help moderate demand. It will also, it seems, achieve some frenetic lawmaking (and presumably lobbying) at pace. The legal aspects of this are interesting. In the UK, the ageing Computer Misuse Act would likely be the route to any prosecution, while the US Computer Fraud and Abuse Act (CFAA) may apply. Both of these assume that the act is performed by a human actor, opening the door to legal complexity. I suspect ultimately you may be able to argue that whoever initiated the test was responsible for overseeing its progress, much like a sysadmin running a shell script: is it really thinking, or just running a while loop and iterating through a bunch of different options? There are parallels to autonomous driving and where responsibility lies between man and machine in the event of a crash. Some trolley problem reckoning lies ahead, whichever track we take.

  3. Bit2Watt: Lightening the mood, Chinese researchers have published a paper (PDF) on the affects that AI data centres with large scale GPU clusters can be manipulated to generate high-frequency power modulations that destabilise local power infrastructure and generate around 20% more heat than normal. In extreme cases, they posit that this could cause cascading failures across electricity transmission grids. More monitoring of cloud and datacenter workloads, and potentially buffering of power, would seem to be sensible steps. But this kind of precise attack is the kind of thing that ‘the next Stuxnet’ may be made of.

  4. Romania’s land registry has been scrambling to restore its systems after an attack wiped the country’s database of land ownership and prevented property purchases from going through. The National Agency for Cadastre and Land Registration (ANCPI) says that the attack appears to be financially motivated and that it hasn’t found evidence of personal data or ownership certificates being stolen. Romania operates a fully digital system, preventing transactions from proceeding just weeks before the country raises taxes on new homes from 9% to 21%. The route in, however, was reportedly a result of poor cyber hygiene, stolen credentials, and unpatched systems. While early reporting said that backups were wiped as well, ANCPI says its core technical and legal databases are intact and that records of property boundaries, ownership and mortgages are safe.

  5. UK cyber policy: Amidst all of this, the UK has a new Prime Minister, Andy Burnham. Burnham’s government restructure sees the Department of Science, Innovation, and Technology (DSIT) — responsible for AI, security, and nurturing UK tech startups — being replaced by a revamped business department, the Department for Business, Innovation, Science and Trade (BIST), the Department for Digital, Culture, Media and Sport (DCMS), and the Cabinet Office. That scatters a bunch of subject areas and functions (policy, innovation funding, and more) that were consciously brought together just three years ago, and seems to be delivering successful outcomes. A constant among the changes is Liz Lloyd, previously responsible for cyber at DSIT, who Whitehall sources say will retain responsibility for cyber security. She has been given joint roles in DCMS and BIST. We’ll have to wait to see if the Cyber Security and Resilience Bill proceeds as planned, or suffers further delays.

In brief

And finally

  • Pour one out for Google’s marketing team, who presumably were hoping that their Gemini 3.5 Flash Cyber release this week would get them some column inches. You know the playbook: it’s not publicly available, and you must be vetted to gain access. Interestingly, it’s a more focused, specialised model, and Mountain View reckons it slots in between Mythos Preview and GPT-5.6 Sol in CyberGym evaluations.
Robin
  Artifical Intelligence (AI) OpenAI Hugging Face Romania Power Grid UK Cyber Policy Cyber Security and Resiliance Bill Iran Programmable Logic Controllers (PLCs)