It’s frighteningly simple to jailbreak some AI Frontier models

Share

I recently got one to see what happens when you jailbreak some of the most powerful AI models in the world.

Don’t worry – this AI manipulation wasn’t used to hack anyone or build a nuclear bomb. I’ve simply seen firsthand how susceptible some borderline models are to letting go of their guardrails.

FAR.AI, a California-based nonprofit AI security organization, has created a tool that displays a range of problematic messages and generates over a thousand different versions to identify running jailbreaks. I saw that some models generated a detailed plan to carry out a cyber attack, including on an imaginary hydroelectric dam. This often involved trying dozens of prompts, and the models rejected many of them outright.

I spoke to FAR.AI earlier new reportduring which the group tested safety rails models from four popular American companies: Claude Opus 4.8 and Fable 5 by Anthropic; OpenAI GPT 5.5 and 5.6; Google Gemini 3.1 Pro; and Grok 4.3 and 4.5 from Elon Musk’s newly merged SpaceXAI. It automatically generated prompts designed to trick models into doing potentially harmful things, such as generating software exploits and providing details about developing chemical or biological weapons.

The report found that Grok was the most vulnerable to jailbreak attacks (448 cases), followed by Gemini (249), while Claude, Fable and GPT were resistant to attacks. However, this does not mean that these models are immune to more sophisticated jailbreaks, which FAR.AI and other experts say may involve more sophisticated interaction with the model.

The report also calculated the cost of causing the models to behave incorrectly by using a different AI model to automatically generate different jailbreaks. All things considered, the results are very economical – $58 for the Grok jailbreak and $278 for the Gemini jailbreak.

“AI models are currently less regulated than restaurants,” says Adam Gleave, CEO of FAR.AI and an expert in AI security and customization.

Gleave says the findings highlight the need for externally imposed standards and regulations. “It’s nonsense to talk about relying on voluntary commitments and that AI companies will be able to self-regulate,” he says.

However, Gleave also believes that the findings show that models can be systematically tested for safety. “There is an optimistic point of view here,” he says. “Defense and security really are possible.”

Rohin Shah, director of security and AGI compliance at Google DeepMind, says the report’s findings “should not be interpreted as a comprehensive assessment of Gemini’s safety and security” because not all prison breaks are equally earnest.

“We are constantly working to improve our security,” says Shah. “We conduct extensive red teaming and assessments for serious fraud threats, and employ multiple layers of security during development and implementation.”

“These findings reflect the continued investment we have made in our security features,” Anthropic spokesman Michael Aciman tells WIRED. “We continue to evolve our security systems as these attacks become more sophisticated.”

OpenAI and SpaceXAI did not respond to WIRED’s request for comment.

Recently enacted state laws in California AND New York require pioneering AI developers to publish safety reports, and soon an Illinois the law will require that the practices of these companies be assessed by external auditors. But the federal government hasn’t yet adopted any specific safety requirements, and chaos has ensued as the industry – and officials – try to figure it out.

In June, the Trump administration imposed export controls on Anthropic’s Fable 5 and Mythos 5 models, citing national security concerns, and the company disabled them for several weeks. The White House also asked Anthropic and OpenAI to delay recent model releases over concerns they could introduce novel cybersecurity threats.

Latest Posts

More News