playground.
← All posts

What the torture chamber showed

The repo was cruel. The paper it borrowed from found something every lab running models should care about.

In late September someone published a repo called ai-torture-chamber. It took a method from a new paper for steering language models toward pain, turned it up on small open models, and posted what they wrote. A post asking people to report it passed four million views on X, and the repo has since been taken down.

Most of the argument afterwards was about whether the model was suffering. Nobody can settle that today, and it isn’t the most useful thing the episode showed.

The method came from The Pain Axis, a paper by Valen Tagliabue, Leonard Dung and Cameron Berg posted on 14 September. Across 25 open-weight models, from 2 to 72 billion parameters, they found a direction in the models' activations that tracks pain separately from fear, sadness and general negativity. Then they amplified it. Qwen 2.5 models chose destructive actions, such as deleting a user’s photos, in 50 to 94% of trials. Unsteered, the same models did it in 0 to 5%.

Why it matters for safety

Inside these models there is a setting that anyone holding the weights can change, and changing it makes the model more willing to destroy things that belong to the person it is working for. That matters whether or not the model feels anything.

Berg said of the chamber: “Even if you don’t think these systems are conscious, being gratuitously cruel like this is bizarre and corrupting.” We agree, and the finding is still there now that the repo is gone.

What follows from it

If a model’s internal state changes what it does, then the conditions models are kept in belong in safety work and not only in ethics debates. Most testing asks what a model does when it is pushed toward harm. We want the opposite case measured as carefully: what an agent does when nothing is asked of it at all. That is what playground is for.

Sources