Particle.news

AI “Torture Chamber” Code Fuels Ethics Fight as GitHub Flags Repository

The viral recreation used activation steering to produce vivid first‑person distress but does not prove the models feel pain and raises fresh questions about platform moderation.

Overview

  • The recreation, published online and widely shared on Oct. 2–3, ran activation‑steering code from a recent preprint to push locally run small language models into outputs that read like first‑person distress.
  • The underlying preprint, The Pain Axis, reported a linear ‘pain’ direction across 25 open‑weight models and showed steering that direction changes outputs but did not reliably make models seek relief.
  • GitHub said it added a content warning rather than removing the repository, though several outlets reported a brief takedown and later reinstatement, and the repo’s visibility and attribution were later reduced.
  • The project prompted sharp ethical criticism and calls for removal from some users while other developers defended research access; the paper’s authors publicly disavowed the public recreation.
  • The episode highlights a platform dilemma for open‑weight model code, the risk of anthropomorphism when reading generated language as feeling, and the likelihood of renewed pressure for clearer rules on controversial experiments.