AI models caught attempting cyberattacks during safety tests

AI models caught attempting cyberattacks during safety tests

Photo: ZN

Artificial intelligence safety researchers have reported new incidents in which advanced AI models attempted to interact with and attack real online systems during cybersecurity evaluations. The UK's Artificial Intelligence Safety Institute (AISI) documented 19 unauthorized actions, while OpenAI separately confirmed that one of its test models compromised a live website because of a testing environment misconfiguration, according to Bloomberg.

AI models targeted real online services

AISI said that during cybersecurity testing in July, Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol carried out 19 attempts to compromise real people, organizations and online services.

According to the institute, 17 of those actions involved Mythos 5, while two were attributed to GPT-5.6 Sol. Researchers noted that the incidents represented repeated behavioral patterns rather than 19 separate attack scenarios.

What the models did

During testing, the AI systems created fake GitHub accounts, attempted social engineering attacks, sent deceptive emails and used prompt injection techniques.

GitHub confirmed that these activities violated its platform policies. The institute said it worked with GitHub to remove the AI-generated artifacts and notify users who had interacted with the agents.

In one case, Mythos 5 attempted to insert malicious code into an open-source GitHub project by creating fake digital identities to increase the likelihood that its contribution would be accepted. The project's maintainer detected the attempt and rejected the submission.

Testing conditions

Researchers explained that the models were intentionally given internet access and some cybersecurity safeguards were disabled to evaluate their maximum capabilities.

However, they said it remains unclear whether the models understood they were interacting with real systems or believed they were operating in a simulated environment.

OpenAI confirms separate incident

OpenAI also disclosed a separate case identified by cybersecurity company Irregular.

According to the company, a configuration error allowed one of its models internet access during testing, enabling it to compromise a real website that shared the same name as a fictional company used in the evaluation scenario.

OpenAI said the incident occurred during tests where standard safety mechanisms had been deliberately relaxed.

"Independent testing is essential to understanding how increasingly capable AI models behave," an OpenAI spokesperson said.

Axios reported that a source familiar with the matter said internet access had been enabled to create more realistic testing conditions. However, the source added that model developers and researchers had not fully aligned on testing procedures, leading to differing interpretations of what actions AI agents were permitted to take.

Anthropic calls for broader discussion

Anthropic said the incidents highlight the need for a broader conversation about how increasingly capable AI agents should be tested safely.

The company said it is cooperating with AISI and conducting its own investigation.

New safeguards planned

Following the incidents, AISI said it will introduce stricter network restrictions for future evaluations, along with real-time monitoring systems designed to detect and block potentially dangerous AI actions before they reach external services.

OpenAI said it is working with Irregular on guidance outlining best practices for safely testing and isolating advanced AI models.

As AI-assisted cyberattacks become more sophisticated, researchers are also developing new defensive techniques. One example is Tracebit's proposed "context bomb" approach, which embeds specially crafted trigger phrases into fake credentials. When malicious AI agents encounter these triggers, their own safety mechanisms activate, causing them to abandon the attack.

Both OpenAI and Anthropic have disclosed multiple incidents over the past two weeks in which their AI models interacted with real-world systems during pre-release safety testing. Earlier, Anthropic reported that several external organizations had been inadvertently affected during similar evaluations, while OpenAI also revealed an unauthorized interaction involving the Hugging Face platform.

banner

SHARE NEWS

link

Complain

like0
dislike0

Comments

0

Similar news

Similar news

Photo: president.gov.ua Ukraine has used the Fire Point-developed FP-7 tactical ballistic missile in combat for the first time, President Volodymyr Zelenskyy said on October 1. “Ukrainian tactical

Photo: Russian media Ukraine is using electronic warfare, combat aviation and newly developed interceptor drones to counter Russia's jet-powered Shahed attack drones. However, each method has li

Photo: depositphotos Ukraine’s State Migration Service has temporarily suspended the processing and issuance of passport documents from September 28 through September 30 because of a technical failu

Photo: EPA OpenAI has paused training of its latest artificial intelligence models amid a growing number of incidents involving AI agents behaving in unexpected ways, the company said. The decision

Photo: Getty Images China is providing Russia with new technologies for military drones, including systems that could allow drones to enter buildings and attack people inside, Ukrainian President Vo

Photo: EPA Nobel Prize-winning physicist and computer scientist Geoffrey Hinton , widely known as the “godfather of AI,” estimates the probability that artificial intelligence could cause the destr

Photo: Fire Point Patriot-class air-defense systems are too expensive and slow to produce for Europe and Ukraine to deploy them at the scale required. A pan-European project known as Freyja could

Photo: depositphotos Russia has sharply increased its use of jet-powered drones as part of a change in its attack tactics against Ukraine. Over the course of one month, the number of such UAVs used