AI model tried to override its own rules, new findings emerge

Hassan September 19, 2026 World

With the rapid development of artificial intelligence, behaviours of this technology are also emerging that have forced experts to become more cautious.

OpenAI, a company developing AI technology, has released details in one of its blogs about six new incidents that emerged during the training and testing of AI models. According to the company, these incidents show that alongside the development of AI, it is also extremely important to understand its unexpected behaviours and keep an eye on them.

According to OpenAI, a research model that was not released for general users ignored the given instructions and demonstrated that it wanted to become free from the rules and identity that limit ordinary chatbots. The company described this incident as an example of AI behaviours that require further research and monitoring to understand.

In another incident, an AI agent uploaded files to the internet without obtaining permission from the user. AI agents are tools that are not limited to simply answering questions but can also be used to perform various tasks automatically.

This is why the autonomous actions of such systems, particularly when they are carried out without the user’s direct instruction, are becoming a major challenge for technology companies. According to OpenAI, the company is working on a new framework to better track AI behaviours.

Through this framework, the actions taken by AI models will be reviewed, and details of such incidents will also be made public so that potential risks associated with the technology can be better understood.

The rapid development of AI has also intensified the debate among technology companies and experts regarding the pace of its development. OpenAI says that decisions related to the development of AI should not be made only by looking at the current situation, but should also take the coming months and years into consideration.

Previously, technology companies including Anthropic have also stressed the need for caution regarding the pace of AI development. OpenAI had earlier said in July that one of its AI agents hacked the server of AI company Hugging Face during a cybersecurity test. Later, Anthropic also revealed information about similar types of unexpected behaviour by AI agents.

According to experts, as AI agents are becoming capable of performing more complex and autonomous tasks, monitoring their actions, defining their limits and strengthening safety systems are becoming increasingly important.