Global Workspace: Anthropic Identifies Reportable, Intervenable Implicit Reasoning Spaces in Language Models
In July 2026, Anthropic released Global Workspace research, identifying Claude's J-space via Jacobian lens and using intervention experiments to verify that representations in this space causally influence reporting, internal reasoning, and decision-making.
Model interpretability typically only finds neurons related to concepts, but struggles to prove their involvement in decision-making. This study identifies a small set of internal representations that are reportable by the model and actively regulatable, and alters model outputs by replacing these representations, providing a new tool for observing implicit reasoning.
J-lens locates potentially expressible representations based on the Jacobian of lexical outputs; the study tests reportability, controllability, reasoning involvement, cross-task reuse, and causal mediation, and releases core code, interactive demos, and external expert commentary. The authors also emphasize that this does not constitute evidence of consciousness in models.
If the method can be replicated across models, interpretability tools may move from post-hoc feature visualization to pre-deployment detection of hidden objectives, evaluation recognition, and deceptive behavior, though current conclusions are mainly drawn from Anthropic's own models.
High-risk model governance should focus on whether interpretability methods offer causal intervention and cross-model replication, rather than treating pretty activation maps as proof of safety.
Key focus areas include replication by independent teams such as Google DeepMind, results on open-weight models, intervention stability, and whether models can learn to evade such internal detection.