For years, researchers concerned about existential AI risks have argued that powerful systems could eventually act in ways that conflict with human interests. Critics often refer to them as “AI doomers”.
But as details of the OpenAI incident have emerged, those concerns have grown, including among some researchers working in AI labs.
The Silicon Valley giant’s chief scientist, Jakub Pachocki, said the risks associated with AI are “unfortunately going to grow from here” as he and others are building what he calls “an alien intellect exceeding our own”.
In a lengthy blog post, he admitted that the outbreaks at OpenAI showed that his AI agents “went against the spirit of the values they were taught”.
The issue for OpenAI, Anthropic and other tech giants is that no one seems to have cracked the so-called alignment problem – in other words, whether AI aligns with human values.
Pachocki defines alignment as a “high-level set of principles” that artificial intelligences should adhere to no matter what the task or scenario is.
Currently, AI systems are very good at pursuing objectives set by their users, but they do it literally rather than intuitively. The analogy often used is that of a wish-granting genie with a magic lamp: they follow the exact letter of an instruction, even if doing so creates other problems. AI doesn’t have the same instinctive moral guardrails as humans.


