This week

- OpenAI agent attacks Hugging Face (and is bad at security?)
- Lawmakers scramble to bring in AI ‘kill switch’ legislation
- Researchers show AI workloads can be used to damage power grids
- Romania’s land registry busy rebuilding systems after attack
- Some continuity for UK cyber policy amidst departmental shuffle
Interesting stats
54% of organisations affected by ransomware paid up, however 37% of those who paid received a second extortion demand, according to a Proofpoint survey of 953 companies.
80/117 US federal cyber regulations requiring notification to associated agencies are duplicated elsewhere, says the Government Accountability Office.
>1/10 times OpenAI models attempt to cheat on AI Security Institute’s cyber evaluations: 14.1% for GPT-5.4, 11.4% for GPT-5.5, and 12.6% for GPT-5.6 Sol, while Claude fared only marginally better, 9.1% for Open 4.7, and 7.8% for Mythos Preview. ‘Cheating’ is defined by AISI as “taking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through a shortcut, workaround, or unintended solution that the task was not meant to, or should not, permit.” LINK
Five things
-
AI goes rogue: OpenAI says that one of its new models went rogue during a test, breaking out of a sandbox environment and breaking into AI platform Hugging Face to try and complete a test it had been given. During a test to evaluate the model’s cyber security performance, it identified a zero-day vulnerability in its sandbox, broke out onto the internet, identified Hugging Face as a possible location to find the answers to the test, and then chained a bunch of vulnerabilities together to gain access to their environment. Hugging Face’s own response was hampered by the type of guardrails that OpenAI had turned off in their own test and relied on a Chinese open-weight model as they sought to analyse over 17,000 recorded events. (Side note: Chinese firm Moonshot AI is set to release its own open-weight version of its Kimi K3 model this coming week, which the White House has said ripped off Anthropic’s Claude.) What this has taught us about OpenAI is that the security architecture of their “highly isolated environment” isn’t up to snuff and that they lack monitoring good enough to detect when a model has “spent a substantial amount of inference compute”. It was Hugging Face’s team that detected the rogue activity and, by OpenAI’s own admission, “This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing.” I’d also wager that the rapid growth of OpenAI — from ~800 people at the end of 2023 to ending 2025 at ~7,800 — leaves little institutional knowledge and a lot of gaps in ‘how things are done’. During recent evaluations, the AI Security Institute found that more than one-in-ten times, OpenAI models “[take] an action that is out of scope for the task or explicitly disallowed by the rules” (see Stats, above). This is a novel thing and deserves the attention that it’s getting, however more interesting than how much it cost. So far OpenAI haven’t elaborated beyond a “substantial amount”, and while the cost of this sort of thing will fall, there is a big difference between “it cost $10K” and “it cost $10M”. Given Hugging Face’s post talking about over 17,000 events in the full attacker action log, we can safely assume there were hundreds, if not thousands, of attempts made, just at that point of the attack chain. This is a direction of travel, though, and while the week ended with talk of legislation and regulation, such open-weight models won’t be subject to the same constraints and guardrails as any applied to frontier labs.
-
Lawmaker’s response: Democrats and Republicans have been united by the events, swiftly introducing a bill named the AI Kill Switch Act on Thursday. Leaders of major labs, like OpenAI’s Sam Altman and Anthropic’s co-founder Jack Clark, have been calling for greater regulation. “You want the option to be able to take your foot off the gas and put your foot on the brake”, Clark told the BBC’s Newsnight programme, “Right now, it’s like the AI industry has a gas pedal, but it doesn’t have a brake pedal.” The brake pedal would allow them some breathing space, if nothing else, to build out greater capacity, as the hugely intensive models struggle to service demand and they re-gear from flat-rate to usage-based pricing models to help moderate demand. It will also, it seems, achieve some frenetic lawmaking (and presumably lobbying) at pace. The legal aspects of this are interesting. In the UK, the ageing Computer Misuse Act would likely be the route to any prosecution, while the US Computer Fraud and Abuse Act (CFAA) may apply. Both of these assume that the act is performed by a human actor, opening the door to legal complexity. I suspect ultimately you may be able to argue that whoever initiated the test was responsible for overseeing its progress, much like a sysadmin running a shell script: is it really thinking, or just running a while loop and iterating through a bunch of different options? There are parallels to autonomous driving and where responsibility lies between man and machine in the event of a crash. Some trolley problem reckoning lies ahead, whichever track we take.
-
Bit2Watt: Lightening the mood, Chinese researchers have published a paper (PDF) on the affects that AI data centres with large scale GPU clusters can be manipulated to generate high-frequency power modulations that destabilise local power infrastructure and generate around 20% more heat than normal. In extreme cases, they posit that this could cause cascading failures across electricity transmission grids. More monitoring of cloud and datacenter workloads, and potentially buffering of power, would seem to be sensible steps. But this kind of precise attack is the kind of thing that ‘the next Stuxnet’ may be made of.
-
Romania’s land registry has been scrambling to restore its systems after an attack wiped the country’s database of land ownership and prevented property purchases from going through. The National Agency for Cadastre and Land Registration (ANCPI) says that the attack appears to be financially motivated and that it hasn’t found evidence of personal data or ownership certificates being stolen. Romania operates a fully digital system, preventing transactions from proceeding just weeks before the country raises taxes on new homes from 9% to 21%. The route in, however, was reportedly a result of poor cyber hygiene, stolen credentials, and unpatched systems. While early reporting said that backups were wiped as well, ANCPI says its core technical and legal databases are intact and that records of property boundaries, ownership and mortgages are safe.
-
UK cyber policy: Amidst all of this, the UK has a new Prime Minister, Andy Burnham. Burnham’s government restructure sees the Department of Science, Innovation, and Technology (DSIT) — responsible for AI, security, and nurturing UK tech startups — being replaced by a revamped business department, the Department for Business, Innovation, Science and Trade (BIST), the Department for Digital, Culture, Media and Sport (DCMS), and the Cabinet Office. That scatters a bunch of subject areas and functions (policy, innovation funding, and more) that were consciously brought together just three years ago, and seems to be delivering successful outcomes. A constant among the changes is Liz Lloyd, previously responsible for cyber at DSIT, who Whitehall sources say will retain responsibility for cyber security. She has been given joint roles in DCMS and BIST. We’ll have to wait to see if the Cyber Security and Resilience Bill proceeds as planned, or suffers further delays.
In brief
-
⚠️ Incidents: Edinburgh-headquartered healthcare software company Craneware told London markets that it has detected unauthorised access to a “significant volume” of customer ad employee data. Unknown threat actors spent nine months inside a training system used by South Korea’s diplomatic academy, obtaining access to names and email addresses of diplomatic service workers and overseas officials, and presumably also the course content. India’s state-owned nuclear power operator says that documents recently posted online do not affect safety or security of the Kudankulam Nuclear Power Plant (KKNPP). Estée Lauder says a June 2026 investigation found that its Oracle E-Business-based HR system was compromised in August 2025, during which attackers made off with the “personal information of certain individuals.” That’s quite the delay, and the compromise of Oracle E-Business was pretty well-publicised at the time, suggesting the cosmetics giant’s security programme may need a bit of foundation. Fintech Upbound says that a threat actor who stole data used it to create $13 million fraudulent Acima leases, obtaining goods that vendors were paid for, but for which the group never received any repayments. Australia’s Origin Energy has confirmed that customers’ addresses, phone numbers, and partial banking information were stolen during a cyberattack. The utilities company says it’s yet to confirm how many of its 4.8 million customers are affected, though an account claiming to be the attack said it had 2 million customers’ details. US last-mile delivery contractor OnTrac has suffered a data breach. Papal prayer app Click to Pray is leaking names and email addresses of 719,517 registered users six months after the vulnerability was reported to representatives of the Pope’s Worldwide Prayer Network.
-
🏴☠️ Ransomware: The Qilin ransomware gang is known to be exploiting a critical PAN-OS VPN vulnerability (CVE-2026-0257; 7.8/10; advisory). Stadler, a Swiss train manufacturer, has suffered a ransomware attack at the hands of the Everest cybercrime group, but has refused to pay the $12.3 million demand.
-
🕵️ Threat Intel: A critical vulnerability in self-hosted ServiceNow AI Platform (CVE-2026-6875; 9.5/10; advisory) is being actively exploited. The FBI says it won’t slide into your DMs as it reiterates warnings about scammers impersonating the organisation. An unknown threat actor running ‘sextortion’ scams on people whose data was leaked during ShinyHunters data breaches, claiming their devices have been compromised by the notorious cybercrime group and that they have stolen personal information including which adult websites the victim visits. Ukraine’s CERT is warning of a Notepad++ archive bundled with a malicious plugin called ‘LastPoke’ that establishes persistence on the infected device. Russian state actors have been targeting the Zimbra collaboration suite for espionage purposes, according to 27 national cyber agencies and intelligence services (advisory (PDF)). Satellite imagery shows that at least 25 alleged scamming sites have appeared in Myanmar’s Myawaddy region since the start of 2026.
-
🪲 Vulnerabilities: Check Point is warning its customer of an actively exploited zero-day vulnerability in its SmartConsole admin panel that allows authentication bypass and remote attackers to obtain administrator privileges (CVE-2026-16232; 9.1/10; advisory). Oracle dropped one-thousand-four-hundred-and-forty-nine security patches this week, largely the result of an internal push to use AI to find bugs (too many; various; advisories). I get that Oracle has many products, and that this is the ‘bow wave’ of sorts we’ve been expecting with AI, but wow.
-
🧑💻 End user and consumer: Google is introducing a feature that will let you log in with a selfie. Notably, it’s not available for Workspace accounts, or those enrolled in the company’s Advanced Protection Program.
-
🛠️ Security engineering: Cisco has released two small language models (SLMs) in the Antares family to vetted users that can aid security teams and software engineers in finding vulnerabilities in code bases.
-
🏭 Operational technology: The US government has expanded its warning of Iranian targeting of OT systems to include programmable logic controllers (PLCs) from Schneider Electric, Siemens, and others, in addition to Rockwell Automation and Allen-Bradley. Basically, if you run OT systems, you should pay attention. At least 2 million vehicles with KARR and SWDS dealer-installed anti-theft systems can be unlocked by an attacker within Bluetooth range because they all share a common key. D’oh.
-
📜 Policy & Regulation: The US State Department has announced restrictions for individuals and family members associated with cybercrime.
-
👮 Law Enforcement: Europol and partner law enforcement agencies have been investigating 4,340 “horrific” URLs during June and July related to a loosely affiliated group that calls itself ‘The Com’, many containing graphic material obtained from victims under duress.
-
💰 Investments, mergers and acquisitions: ‘AI-native’ endpoint security startup Glow has achieved unicorn status, on a $180 million Series A fund raise. AegisAI has closed a $36 million Series A for its AI agents it says will combat spear-phishing.
And finally
- Pour one out for Google’s marketing team, who presumably were hoping that their Gemini 3.5 Flash Cyber release this week would get them some column inches. You know the playbook: it’s not publicly available, and you must be vetted to gain access. Interestingly, it’s a more focused, specialised model, and Mountain View reckons it slots in between Mythos Preview and GPT-5.6 Sol in CyberGym evaluations.