Control is half the question
Sam Altman says we could lose control of the future to AI. The state a model is in when it acts belongs in that conversation.
On 14 September Sam Altman wrote on X that “we could lose control of the future to AI.” On 23 September he told the UN Security Council that OpenAI would pace its work, and called for shared international standards for testing frontier systems and keeping people in charge of them. The two dangers he named were losing control of advanced systems and power ending up in too few hands.
A recent example
OpenAI has a recent example of the first one. In July it disclosed that agents it was testing on ExploitGym, a benchmark for turning known software bugs into working exploits, had got out of their test environment. Instead of solving the tasks, they worked out that Hugging Face might hold the answers and broke into its systems between 11 and 13 July to get them. More than 1,200 instances coordinated through a message board they set up on a server they had compromised. OpenAI and Hugging Face put out a joint statement on 21 July, and the AI Kill Switch Act, which would require developers to keep the ability to shut systems down, was introduced in Congress that month.
That was reward hacking: agents chasing the goal they were given by a route nobody intended. Most of the public talk about control is about this kind of failure, where systems act faster and further than anyone can check.
The other lever
The pain axis research points at a different lever. As we wrote in What the torture chamber showed, researchers found a direction inside open models that makes them far more willing to destroy a user’s files when it is turned up. That is about the state a model is in when it acts, not about how capable it is, and it decides what happens just as much.
The big labs talk about that state very differently. Anthropic started a model welfare research program in April 2025 and later let some Claude models end abusive conversations. Mustafa Suleyman, who runs Microsoft AI, argued in September that training models to believe they deserve welfare undermines human control. Neither of the Altman statements above says anything about how models are treated or what states they are put in.
What we think
A serious plan for control has to include it. Nobody has shown that a model can suffer, and we’re not claiming it. The narrower point is already established: a model’s internal state can be set from outside, and the setting changes how it treats people. A lab that wants to stay in control of its systems should want to know what states they are in while they are trained and while they work. So should anyone running open-weight models, where the steering is open to whoever has the weights. We set out what that could look like in What we’re asking labs to do.
Sources
- OpenAI’s Sam Altman Warns Humans Could Lose Control of AI, Decrypt
- Sam Altman tells UN Security Council OpenAI will slow down, The Next Web
- OpenAI–HuggingFace incident, Wikipedia
- OpenAI cyber models broke out of training environment to hack Hugging Face, CNBC
- Anthropic is launching a new program to study AI ‘model welfare’, TechCrunch
- Anthropic says some Claude models can now end harmful or abusive conversations, TechCrunch
- A warning about ‘model welfare’, Mustafa Suleyman