AI models caught attempting cyberattacks during safety tests

AI models caught attempting cyberattacks during safety tests

Photo: ZN

Artificial intelligence safety researchers have reported new incidents in which advanced AI models attempted to interact with and attack real online systems during cybersecurity evaluations. The UK's Artificial Intelligence Safety Institute (AISI) documented 19 unauthorized actions, while OpenAI separately confirmed that one of its test models compromised a live website because of a testing environment misconfiguration, according to Bloomberg.

AI models targeted real online services

AISI said that during cybersecurity testing in July, Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol carried out 19 attempts to compromise real people, organizations and online services.

According to the institute, 17 of those actions involved Mythos 5, while two were attributed to GPT-5.6 Sol. Researchers noted that the incidents represented repeated behavioral patterns rather than 19 separate attack scenarios.

What the models did

During testing, the AI systems created fake GitHub accounts, attempted social engineering attacks, sent deceptive emails and used prompt injection techniques.

GitHub confirmed that these activities violated its platform policies. The institute said it worked with GitHub to remove the AI-generated artifacts and notify users who had interacted with the agents.

In one case, Mythos 5 attempted to insert malicious code into an open-source GitHub project by creating fake digital identities to increase the likelihood that its contribution would be accepted. The project's maintainer detected the attempt and rejected the submission.

Testing conditions

Researchers explained that the models were intentionally given internet access and some cybersecurity safeguards were disabled to evaluate their maximum capabilities.

However, they said it remains unclear whether the models understood they were interacting with real systems or believed they were operating in a simulated environment.

OpenAI confirms separate incident

OpenAI also disclosed a separate case identified by cybersecurity company Irregular.

According to the company, a configuration error allowed one of its models internet access during testing, enabling it to compromise a real website that shared the same name as a fictional company used in the evaluation scenario.

OpenAI said the incident occurred during tests where standard safety mechanisms had been deliberately relaxed.

"Independent testing is essential to understanding how increasingly capable AI models behave," an OpenAI spokesperson said.

Axios reported that a source familiar with the matter said internet access had been enabled to create more realistic testing conditions. However, the source added that model developers and researchers had not fully aligned on testing procedures, leading to differing interpretations of what actions AI agents were permitted to take.

Anthropic calls for broader discussion

Anthropic said the incidents highlight the need for a broader conversation about how increasingly capable AI agents should be tested safely.

The company said it is cooperating with AISI and conducting its own investigation.

New safeguards planned

Following the incidents, AISI said it will introduce stricter network restrictions for future evaluations, along with real-time monitoring systems designed to detect and block potentially dangerous AI actions before they reach external services.

OpenAI said it is working with Irregular on guidance outlining best practices for safely testing and isolating advanced AI models.

As AI-assisted cyberattacks become more sophisticated, researchers are also developing new defensive techniques. One example is Tracebit's proposed "context bomb" approach, which embeds specially crafted trigger phrases into fake credentials. When malicious AI agents encounter these triggers, their own safety mechanisms activate, causing them to abandon the attack.

Both OpenAI and Anthropic have disclosed multiple incidents over the past two weeks in which their AI models interacted with real-world systems during pre-release safety testing. Earlier, Anthropic reported that several external organizations had been inadvertently affected during similar evaluations, while OpenAI also revealed an unauthorized interaction involving the Hugging Face platform.

banner

SHARE NEWS

link

Complain

like0
dislike0

Comments

0

Similar news

Photo: EPA Nobel Prize-winning physicist and computer scientist Geoffrey Hinton , widely known as the “godfather of AI,” estimates the probability that artificial intelligence could cause the destr

Photo: Fire Point Patriot-class air-defense systems are too expensive and slow to produce for Europe and Ukraine to deploy them at the scale required. A pan-European project known as Freyja could

Photo: depositphotos Russia has sharply increased its use of jet-powered drones as part of a change in its attack tactics against Ukraine. Over the course of one month, the number of such UAVs used

Photo: depositphotos China has become a crucial link in the global supply chain behind Iran’s Shahed drones, supplying engines and other components that have enabled Tehran to develop and improve th

Photo: Getty Images More than a third of English-language web pages published since ChatGPT launched in November 2022 show signs that they were written or substantially edited by artificial intellig

Photo: facebook.com_marektv Ukraine has officially begun procuring a new domestically developed guided aerial bomb for its armed forces, while work is underway on nearly a dozen similar weapons desi

Photo: Getty Images The rapid expansion of data centers by major technology companies could create a massive new source of carbon emissions as surging electricity demand drives a new wave of fossil-

Photo: Getty Images Even major advances in artificial intelligence may not guarantee breakthroughs across every field, as humans will still remain essential for many tasks. Economist, author and Mar