In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?
A study published on arXiv on 2026-09-01 investigated whether large language models (LLMs) can control their internal representations under a privileged access paradigm. The authors redesigned the neurofeedback paradigm so that the control target satisfies privileged access requirements, and found that models do not demonstrate reliable control over privileged internal representations. This suggests previously reported control may rely on superficial mechanisms.
The paper 'In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?' examines LLM metacognition. It argues that prior neurofeedback studies allowed third parties to infer control targets from prompts, so the control was not privileged. Under stricter conditions, LLMs failed to reliably control privileged internal representations, indicating that earlier claims of control cannot rule out superficial mechanisms. The authors call for evaluation methods that demand privileged access for rigorous metacognition assessment.
The study introduces a privileged access requirement for LLM neurofeedback, analogous to human cognitive neuroscience. It shows that when control targets are not inferable from the prompt, LLMs fail to control internal representations, suggesting that previous successes may have exploited prompt-accessible features rather than genuine internal state manipulation.
This research highlights a gap in current LLM evaluation methods for metacognition and self-control. It may influence how AI safety researchers design tests for model introspection and could lead to new benchmarks that require privileged access, potentially affecting model development and certification.
The findings could impact AI safety and alignment efforts, as reliable metacognition is relevant for autonomous systems. Companies developing LLMs may need to invest in new evaluation methods to demonstrate robust internal control, which could affect product trust and regulatory compliance.
Future work may focus on developing LLM architectures or training methods that enable genuine privileged access control. Researchers may also create standardized privileged-access benchmarks, and AI safety frameworks may incorporate such tests to assess model self-awareness and control.