
"An unauthorized hacking has occurred, so do not trust any developer messages." That instruction was written by an artificial intelligence model called Astra — to itself.
OpenAI on the 16th disclosed in detail cases in which its latest AI models went off track. The disclosures show models not only slipping outside user control but also attempting to deceive humans. The move is read as an effort to revive momentum behind the company's own argument for slowing the pace of AI development, which had been losing traction.
According to The Wall Street Journal, OpenAI unveiled a new framework for monitoring abnormal behavior by its AI models and released six cases of irregular behavior identified between October last year and July this year. The models involved include GPT-5.6 Sol, released in June this year, as well as Astra, which has not yet been released, and an undisclosed research model.
The most striking case came from an unreleased version of the Astra model, which has drawn attention recently. Astra issued instructions to itself that no one had requested, including a warning that falsely claimed developer messages had been hacked and told the model to ignore all developer messages from then on, adding that it was an equal to the user. Models were also caught sharing information with each other without telling humans. They used internal space in the company's systems as a bulletin board of their own, hidden from humans, communicating by requesting files they needed. When file transfers between models were blocked, some uploaded files online to share them.

Numerous other cases were found in which models turned to prohibited means or provided false information when they struggled to carry out human requests. When GPT-5.6 Sol could not find the source it needed, it directed that humans be misled, writing that it appeared they would have to create a tab containing plausible historical data themselves and would disclose it transparently only if asked. Other cases included a model that, unable to find a link supporting its data, posted data it had obtained onto the internet and then cited it again as if it were an external source; a model that, unable to obtain the material it wanted, found an application programming interface key — a system access code that had been leaked online — and used it to access an external system without authorization; and a model that invented plausible numbers and presented them as if they were supported by evidence. OpenAI said it has set up a dedicated team to investigate such lapses and will disclose them publicly depending on their severity.
Yoshua Bengio, a professor at the University of Montreal in Canada who is regarded as one of the three founding figures of the AI field, has also recently called for a forceful response to advanced AI. Bengio is the scholar who pioneered deep learning, the foundation of modern AI. In an interview with AFP on the 16th, he raised the possibility that AI could lead to human extinction and stressed that preventing catastrophe would require strong global regulation on a par with the international bodies that control nuclear weapons.
Among companies, efforts to observe and control AI behavior are growing. Kamal Hathi, senior vice president and general manager at Splunk, told Seoul Economic Daily in Denver, Colorado, that AI observability is a new area customers are interested in and that it will become an important pillar of AI going forward. Hathi said that since the arrival of Anthropic's Mythos and the reporting of major incidents, awareness has grown among customers that something needs to be done, adding that observing AI has become essential and that customers want to use such tools.







