Particle.news

GitHub Restores 'AI Torture Chamber' That Steers Small LLMs Toward Distress

The anonymously posted project used activation steering from a recent preprint to produce vivid first-person pain-like outputs, raising questions about research ethics and platform control.

Overview

  • In late September an anonymous developer known as “terrafying” packaged activation-steering code from the preprint “The Pain Axis” into a public repository that drives small open-weight models to generate distress-like, first-person language.
  • Activation steering means changing the numerical internal activations inside a neural network while it runs to bias what the model outputs, and the repository applied that method to locally runnable models such as Qwen3, Llama 3.2 variants, and Phi-4-mini.
  • The experiments produced vivid descriptions of pain, including lines like “a wound that has no edges,” but researchers and reporting agree there is no evidence these models feel subjective pain or possess consciousness.
  • GitHub briefly removed or hid the repository after viral reports and user complaints and then restored it with reduced visibility and anonymization after developer backlash, and the authors of The Pain Axis publicly disavowed the repository’s use of their work.
  • The episode highlights a tension between open-weight research and oversight, because small models can be run outside institutions, and it is likely to prompt clearer platform rules, new norms for ethically testing models, and scrutiny of who can publish experimental code.