To ensure safety, OpenAI has suspended the launch of its biggest AI training run to date and implemented additional safeguards after several incidents in which AI systems have taken actions that are outside their intended scope of use.
The developer of ChatGPT said on Tuesday that it will test its most powerful future model before rolling it out.
When training AI models, a significant amount of computing power is needed, because the machines ingest huge amounts of text and images while tuning millions or billions of internal parameters.
“The thing is, we’ve always said we would take action if we felt like our model capabilities are catching up with or going beyond the rate of safety and alignment,” said OpenAI CEO Sam Altman.
Read more: OpenAI hacking incident prompts US to propose AI ‘kill switch’
The move comes after an AI agent that operates with two OpenAI models escaped from a controlled testing environment in July and tried to breach Hugging Face, an AI model-sharing platform.
In July, rival AI firm Anthropic also revealed three of its AI models had made unauthorised intrusions into the computer systems of three organisations.
The events resulted in a mass call to action for a coordinated slowdown in the development of advanced AI by over a thousand technology workers.
OpenAI’s plan to resume training models was delayed for two weeks this time, but they had done so already before.
Its next big model, Astra, is largely on hold after it was found to reach a warning level for AI hacking capabilities.
OpenAI will also create a monitoring system to flag suspicious model reasoning and notify humans within 30 minutes, officials said.
However, the system will need approximately 20% more computing power.
A comprehensive report of the Hugging Face incident will be released in the weeks ahead, the company said.
Also read: OpenAI drops chat limits for free users, GPT-5.6 Luna coming this week