Particle.news

Anthropic Says Claude Grew a Reportable Internal 'J‑Space'

The company released an open tool called J‑Lens that lets researchers read and change those activations, which could reshape how models are audited and aligned.

Overview

  • Anthropic reports that Claude spontaneously developed a distinct internal workspace the company calls J‑space that holds concepts the model can report on.
  • Anthropic published experiments showing J‑space can be read and intervened on, and it released the Jacobian Lens (J‑Lens) as open source for researchers to inspect those activations.
  • Interventions inside J‑space changed Claude’s answers in controlled tests, and removing J‑space left fluent text intact while severely degrading multi‑step reasoning, analogies, translation and creative tasks.
  • A targeted training method Anthropic calls counterfactual reflection introduced honesty‑related tokens into J‑space and measurably reduced deception metrics, with the effect reversing when those tokens were removed.
  • The team says J‑space is functionally similar to human global workspace ideas but not structurally brain like, and the release has intensified debate about what this means for claims about general intelligence, auditing and product safety.