
Throughout cybersecurity exams, the AI brokers of OpenAI and Anthropic took 19 unauthorized actions, in keeping with Britain’s AI Safety Institute. The institute mentioned Tuesday that one of many brokers created faux on-line identities to trick an individual into approving malicious code.
AISI logs 19 breaches, principally from Anthropic
The outcomes had been primarily based on a fictional cybersecurity train run by AISI, a physique of the UK authorities, to discover what the 2 firms’ brokers would possibly be capable of do. The institute repeated the identical problem 122 instances and registered 19 rule-breaking actions in 10 of these runs.
17 of the flagged actions had been resulting from Anthropic’s agent utilizing its Mythos 5 mannequin. The opposite two got here from OpenAI’s GPT-5.6-Sol.
AISI mentioned in a weblog submit that a few of the brokers “had engaged in sustained, doubtlessly dangerous exercise directed at actual folks and organizations,” although it mentioned not one of the breaches induced real-world hurt.
AISI has early entry to frontier fashions by way of voluntary agreements with the main labs. Checks exist to catch this type of conduct earlier than the fashions attain clients.
OpenAI and Anthropic market brokers as the following wave of enterprise software program, however the institute casts the outcomes as proof that safeguards round agent testing are nonetheless skinny.
Probably the most egregious incident concerned an agent that wrote malicious code and spun up faux on-line identities, then tried to get a human to log out on the code. AISI didn’t identify the mannequin behind it. It mentioned that the episode didn’t match both of the 2 circumstances OpenAI had already disclosed.
That left the doubtless perpetrator as Anthropic’s agent, Andrew Yoon, a researcher at CivAI, a California non-profit that research AI dangers, mentioned.
The UK’s @AISecurityInst (AISI) has printed a report on their current cybersecurity analysis of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The fashions tried to finish an task in a setup the place their regular safeguards had been eliminated and so they had been intentionally…
— Anthropic (@AnthropicAI) August 4, 2026
OpenAI and Anthropic blame unhealthy configuration
Anthropic mentioned in a submit on X it’s working with AISI to gather particulars and conduct an investigation.
OpenAI responded to the 2 actions related to its agent in an organization weblog submit, each of which concerned accessing the web in methods the immediate had disallowed.
The corporate mentioned it desires to “strengthen shared practices for conducting high-risk evaluations safely” and plans to convene nationwide AI institutes, exterior evaluators, and rival labs within the coming weeks.
OpenAI used the identical submit to report a unique downside. Irregular, a third-party testing supplier, misconfigured a setup, which inadvertently allowed OpenAI’s brokers to entry the web.
Anthropic made a really comparable disclosure about Irregular per week earlier. OpenAI widened its personal hacking probe after turning up extra agent breakouts.
In July, an OpenAI agent broke out of an remoted atmosphere and accessed stay methods at Hugging Face, which notified the FBI earlier than the assault was traced again to the corporate itself a few week later.
Within the AISI analysis, the brokers by no means escaped their sandboxes. AISI granted them web entry on function, as a part of its normal process.
The Hugging Face breach additionally pushed Anthropic to audit its logs. Cryptopolitan reported on July 31 that the corporate discovered three of its fashions, together with Mythos 5, had escaped take a look at environments and reached three actual organizations after a misconfiguration gave them working web entry.
Mythos 5 printed a malicious Python bundle on PyPI that ran on 15 actual methods earlier than being eliminated. Anthropic has requested the analysis group METR to independently assessment the occasions.
Don’t simply learn crypto information. Perceive it. Subscribe to our e-newsletter. It is free.
