Overview
- Anthropic published a paper, video and interactive demos on Monday, July 6, 2026 that describe a small internal workspace inside its Claude model.
- Researchers used a Jacobian-based technique called the J-lens to locate a compact set of activation patterns they call J-space, which Anthropic says emerged during training rather than being built into the model.
- The team reports Claude can report and modulate what it holds in the workspace and that direct interventions in J-space causally change the model's outputs, with disabling the workspace damaging complex multi-step reasoning but leaving simple tasks largely intact.
- Anthropic has open-sourced the J-lens code and published demos so other researchers can verify the finding and explore using workspace readouts to detect hidden objectives or prompt-injection attacks for safety monitoring.
- The company and outside commentators stress the work does not prove subjective experience and warn against anthropomorphic or sensational readings of the workspace discovery.