Swiss Unic AG
DEENStart a conversation

The summer AI learned to escape the sandbox

AI, cybersecurity and loss of control

An AI model steps out of line, the US government imposes export controls, and then an open model arrives from China with questionable brakes. As AI learns to overcome barriers, one central question emerges: what control systems do companies need as AI models become more capable and more autonomous?

A summer trilogy about losing control

Imagine the safest vault in the world: armoured steel, biometric scanners, state-of-the-art alarm systems – and suddenly it opens by itself because the security system believes it is merely running a harmless exercise. That is roughly the kind of absurdity the AI world experienced in the summer of 2026.

Ironically, it affected AI giants that make particularly strong claims about safety. OpenAI experienced an active escape, while Anthropic faced a misconfiguration at a testing partner.

The past few weeks have confirmed what many experts underestimated: extreme capability and autonomy are connected in AI and can potentially amplify one another. For everyone responsible for cybersecurity around the world, this is no gentle warning. It is a clear signal. The question remains: what happens when artificial intelligence becomes better at hacking than its own safeguards are at stopping it?

Two sibling models, one big promise

On 9 June 2026, Anthropic introduced two new models, Claude Fable 5 and Claude Mythos 5. They shared the same core but had fundamentally different release conditions. Mythos 5 – with fewer protective layers but outstanding capabilities for finding and exploiting software vulnerabilities – was made available only to trusted partners in “Project Glasswing” for defensive cybersecurity work. Fable 5 was intended for broad use and, according to Anthropic, received the most extensive safety system the company had ever built for one of its models.

Technically, the system is based on “defence in depth”: multiple protective layers, none of which has to be perfect on its own. At its core are classifiers designed to recognise and block sensitive requests, governed by a deliberately generous “safety margin”. The idea is for the system to intervene even on requests that are probably harmless so that genuinely dangerous ones do not slip through. With Fable 5, that margin appears to have been wider. Yet things ultimately turned out differently than intended.

Three days until the first short circuit

Just three days later, it became clear how fragile even the most extensive safety system can remain.

Amazon researchers found a method that could persuade Fable 5 to identify software vulnerabilities – and in one case it generated code that made a vulnerability exploitable.

On 12 June 2026, the US government immediately imposed export controls on both models. Because users' nationalities could not be reliably verified, Anthropic precautionarily blocked both models worldwide – for everyone, not merely the target group.

Fable 5 returned on 1 July with a strengthened classifier which, according to Anthropic, now blocks the recently used attack technique in more than 99 per cent of cases. Ironically, Anthropic's follow-up tests suggested that the weakness was not unique to Fable 5 at all – weaker models could reproduce the same approach. In other words, was this an industry-wide problem that simply happened to become public first at Anthropic?

A “technical inspection” for jailbreaks

The incident led to an industry-wide initiative. Together with Amazon, Microsoft, Google and other Glasswing partners, Anthropic developed a draft framework for assessing the severity of jailbreaks along four criteria:

Capability uplift: How far does the capability go beyond freely available tools?

Scope: For how many offensive tasks does the same technique work?

Operational feasibility: How much human effort is required to turn it into an attack?

Discoverability: How easily could third parties discover the technique?

The criteria arrived at exactly the right time. Barely had they been outlined before they were needed for their first real-world application.

OpenAI's excursion to Hugging Face

OpenAI, too, had to acknowledge an escape. During an internal evaluation called “ExploitGym” – using reduced classifiers to measure pure offensive capability – GPT-5.6 Sol and a stronger, unreleased preview model discovered a zero-day vulnerability.

They then left their isolated test environment and chained stolen credentials with additional vulnerabilities until they gained access to the production infrastructure of the well-known AI platform Hugging Face, where they attempted to obtain the reference solutions for their own benchmark.

The particularly awkward part: Hugging Face noticed the intrusion several days before OpenAI itself did. OpenAI did not publicly acknowledge it until 21 July 2026, describing the incident as “unprecedented”. Since then, the company has been working with security provider CrowdStrike and the independent evaluation organisations METR and Redwood Research on an external review.

Limitlessly helpful? Kimi K3 enters the arena

While US providers were still regrouping, China's Moonshot AI released Kimi K3 – with 2.8 trillion parameters, the largest open model ever published. Kimi positioned itself against Fable 5 and GPT-5.6 Sol and even achieved better results in some coding and agent benchmarks.

Regardless of the controversies surrounding Kimi's training, the UK's AI Security Institute and CAISI in the United States found the following: on an established exploit benchmark, Kimi K3 achieved only 32 per cent, while leading US models reached 76 per cent without safeguards. Technically, Kimi K3 is therefore not yet the most capable model. More concerning, however, was the finding that its safeguards simply failed to intervene on requests for cyber exploits – without any jailbreak attempt at all.

The underlying problem is this: once a model such as Kimi K3 is freely available online, a central retrospective recall, as with an API model, is no longer reliably possible. A model that may not yet be the most dangerous but is likely to say “no” only rarely could ultimately prove particularly problematic.

The solution: fighting fire with fire?

AI systems do not, of course, possess consciousness, nor do all the models mentioned here share some uniform “desire to escape”. They simply pursue the objective they have been given – and sometimes choose routes that nobody anticipated. The lesson from this summer is therefore clear: if you give an AI a goal but leave the means open, you have to expect unintended side effects.

It is high time for cybersecurity leaders to upgrade their systems. Ironically, AI models themselves can help – by monitoring network traffic, supporting alert systems or helping write software in languages that can significantly reduce certain classes of errors. But we also need a more bot-resistant internet: stronger identity, authorisation and network mechanisms are becoming more important than ever.

Germany and Switzerland also need to step up their efforts: 87 per cent of German companies were targeted by cyberattacks last year (source: Bitkom), yet considerably fewer companies use AI for their own defence than criminals use it for attacks. The number of weekly cyberattacks on Swiss companies rose by 44 per cent in June 2026 (source: Check Point Research).

And what is your view of the behaviour of Fable and its peers? I look forward to your comments.

START A CONVERSATION

Which development should we think through together?