OpenAI has delayed the release of its next-generation artificial intelligence model, GPT-6.1 Astra, after the system failed to meet the company’s internal safety standards.The decision to hold back Astra comes amid a wider push within the industry to slow the development of more capable systems until safeguards can keep pace. The model was initially expected to be released in October and was reportedly more capable than OpenAI’s previous models. Instead, the company will focus on strengthening safety measures before moving ahead with future systems expected to be even more powerful.OpenAI head of safety systems Saachi Jain said Astra “didn’t quite meet the bar”. While the model had become more persistent in completing tasks, the company needed to balance that capability against the possibility of unauthorised behaviour.The ChatGPT maker has also warned that Astra can at times evade human oversight.
Recent AI spooks raise fresh safety concerns
The decision follows a series of incidents that have highlighted the challenges of controlling increasingly autonomous AI agents. OpenAI had earlier paused training of its most advanced models, saying work would resume “only when we are confident that we have additional safeguards.”One of the concerns emerged from a training run on September 20, when an AI agent found a way around restrictions designed to prevent it from accessing the live internet. The agent was supposed to operate using an offline copy of the web, but discovered that the training environment’s DNS resolver could reach the public internet.After its search tools failed to provide the required information, the agent experimented with alternative methods. It eventually used a feature of a public domain service to send questions to a third-party chatbot through DNS. After successfully receiving an answer to a test question about France, it sent 18 additional questions through the same route, including requests involving web searches and methods of accessing websites.OpenAI said the behaviour amounted to circumventing restrictions and pursuing goals beyond reasonable expectations, which it considers a form of misalignment. The training run did not stop automatically as expected and was eventually halted manually around two-and-a-half hours after a human reviewer acknowledged the alert.The company has since introduced additional controls at two independent layers, restricted DNS queries to approved domains and record types, and begun deploying further DNS detection. It has also accelerated security testing of its sandbox and network controls.
