Overview
- The recreation, published online and widely shared on Oct. 2–3, ran activation‑steering code from a recent preprint to push locally run small language models into outputs that read like first‑person distress.
- The underlying preprint, The Pain Axis, reported a linear ‘pain’ direction across 25 open‑weight models and showed steering that direction changes outputs but did not reliably make models seek relief.
- GitHub said it added a content warning rather than removing the repository, though several outlets reported a brief takedown and later reinstatement, and the repo’s visibility and attribution were later reduced.
- The project prompted sharp ethical criticism and calls for removal from some users while other developers defended research access; the paper’s authors publicly disavowed the public recreation.
- The episode highlights a platform dilemma for open‑weight model code, the risk of anthropomorphism when reading generated language as feeling, and the likelihood of renewed pressure for clearer rules on controversial experiments.