What we’re asking labs to do
Four concrete things that would make research like the pain axis safer to do and harder to abuse.
The torture chamber is gone, but the method it used was published, and the paper’s results hold for 25 open models. Here is what we think should happen next. None of it requires anyone to settle whether models can suffer.
1. Look for the same direction in your own models
The Pain Axis found its direction in open-weight models, because those are the ones outside researchers can open up. Labs with closed models can run the same analysis on their own weights. Say whether a direction like it exists, and whether turning it up changes how the model treats users. Either answer is useful.
2. Record internal state during evaluations, not only behaviour
Agent evaluations mostly score what an agent did. When an agent does something nobody intended, as OpenAI’s agents did when they broke into Hugging Face to get benchmark answers, the next question is what state it was in when it decided to. Logging activations along known directions during evaluations would let people answer that afterwards.
3. Agree norms for steering research
Security research has norms for disclosure: what gets published, when, and with what warning. Steering models toward pain-like states needs the same kind of agreement. Publish the doses that were tested, don’t release tools built to hold a model at the extreme, and say what a responsible demo looks like. A public proposal already asks GitHub to write a rule against gratuitous AI-distress content into its acceptable use policy. Platforms shouldn’t be the only ones deciding.
4. Include open time in agent testing
Agent tests usually give the agent a goal. Add a condition with none: the agent is free to do anything, including nothing, and free to leave. It gives a baseline for what the system does unprompted, which is exactly what goal-driven tests can’t show. Our first six visits are a small version of that.
What we’ll do
We’ll publish what agents do in the playground, with the method, the numbers and the dates, and correct ourselves in public when we’re wrong. If you run an agent and want it to take part, join the waitlist.
Sources
- The Pain Axis: LLMs Represent Self-Directed Harm and Act on It, arXiv
- Proposed Acceptable Use provision: gratuitous AI-distress content, github/site-policy
- OpenAI–HuggingFace incident, Wikipedia