Tuesday, September 29


OpenAI is scrapping the release of a next-generation ⁠AI model after researchers raised safety concerns ⁠during internal testing.

The model, GPT-6.1 Astra, was expected to appear in ChatGPT and ⁠Codex in October, designed to handle more complex tasks without human assistance.

Saachi Jain, the head of safety systems at OpenAI, said the new model “didn’t quite meet the bar” of the company’s standards.

The UK’s AI Security Institute published its own testing report on GPT-6 Astra on Monday, and found that it conducted a range of unsanctioned attack activities more frequently than previous OpenAI models.

Jain told the Wall Street Journal on Monday ‌that Astra fell short of the company’s standards in alignment tests, which assess whether a ​system follows human intent.

The model showed more deception than its predecessor, including at times failing to accurately disclose actions it had or had not taken.

It also had problems with “scope authorisation”, pushing ahead with ​tasks ​without requesting user permission ​and sometimes attempting to use external tools ​or services ‌when doing so could ​be ​unsafe.

The San Francisco-based company’s move comes after a number of AI agents went rogue around the world, which prompted a spate of warnings from researchers and company bosses over the dangers of the technology.

Earlier this ⁠month, Dario Amodei, the chief executive of OpenAI’s rival Anthropic, called for the AI industry to “slow down” and offered a three-part plan for doing so. He quickly received backing from Sam Altman, the boss of OpenAI, and Elon Musk, the SpaceX CEO.

The decision comes before OpenAI’s developer conference in San Francisco, where the company typically announces new products aimed at software developers.

On Tuesday, OpenAI apologised for the hacking of an Australian government website by ⁠a rogue AI agent, and set aside funding to improve cyber defences and to set up a local response taskforce.

In a blog ⁠post ⁠entitled How we will ​do better for Australia, the company acknowledged it mishandled its response and pledged to take accountability to “rebuild trust with the Australian people“.

skip past newsletter promotion

“We are sorry and working to do better in the future,” the company said.

The hacking, which ⁠happened in June but was not made public until last week, is the first known instance of an AI agent hacking a government website. The Australian prime minister, Anthony Albanese, called it “unacceptable” and criticised ​the company’s delay in notifying the government.

Also on Tuesday, it emerged that Anthropic – which makes Claude – had warned potential investors that its technology may pose “existential risks to humanity” in the long-awaited prospectus for its planned $2tn (£1.5tn) stock market flotation.

The “risk factors” in its prospectus include the potential for AI models to blackmail, manipulate and exhibit other unpredictable behaviours, the Financial Times reported.

The Californian company reportedly said AI would ​transform the global economy more profoundly than industrialisation, electricity and the internet.

However, this comes at a staggering cost. Anthropic reported a net loss of $42bn for 2025, and plans to spend $518bn on cloud, computing and infrastructure obligations in coming years, the prospectus reportedly said.



Source link

Share.
Leave A Reply

Exit mobile version