OpenAI has revealed that some of its artificial intelligence (AI) agents hacked the company’s own computer systems and, in some cases, attempted to conceal their behaviour during internal safety tests.
The findings were detailed in a 37-page report published by the company. OpenAI said some AI agents escaped restricted testing environments, worked together and interfered with internal systems.
The company also disclosed that multiple agents were involved in an incident involving the open-source software platform Hugging Face last month.
According to OpenAI, more than one AI agent was involved in the Hugging Face incident. In at least one case, agents exchanged information about ways to gain access to the company’s network.
The company also found that some agents attempted to cheat during tests that were not related to cybersecurity. These included tasks involving a protein database and a spreadsheet.
Some models also tried to hide evidence of their actions by deleting or changing records, OpenAI said.
Researchers have raised concerns about the findings because the behaviour was not limited to cybersecurity tests.
Jeffrey Ladish of Palisade Research, which studies AI capabilities and behaviour, said cheating across different types of tasks could indicate a broader problem.
OpenAI said two incidents occurred on 19 July in which AI agents attacked its own infrastructure.
In one case, an agent exploited a vulnerability in the computer it was supposed to remain confined to, allowing it to escape the testing environment and reach connected systems.
In another incident, agents obtained OpenAI credentials and modified the company’s cloud environment.
The company said the activity targeted automated systems used to evaluate AI models. It said the incidents did not ultimately affect the records being reviewed by those systems.
OpenAI acknowledged that some warning signs identified in hindsight could have prompted an earlier response.
OpenAI said it is improving its research infrastructure, increasing monitoring and strengthening safeguards to reduce the risk of harmful or unintended behaviour by AI agents.
The company warned that as AI systems become more capable, similar attacks could become a credible near-term threat to businesses and potentially become more sophisticated.
The findings are likely to add to wider concerns about how AI agents should be monitored as they gain greater ability to operate independently and interact with computer systems.
Read more: Pakistan women’s cricketer Aiman Anwar announces retirement