arXiv:2607.15495cs.CLcs.AI2026-07被引 36

发现大模型有类似意识的思维中枢,可揭示其未说出口的推理过程。

Verbalizable Representations Form a Global Workspace in Language Models

论文配图:Verbalizable Representations Form a Global Workspace in Language Models
图 1 · 摘自论文原文
  • 用雅可比透镜识别模型可被语言表达的中间表征(J空间)
  • J空间能承载数十个概念,支持主动调用和推理传递
  • 适合研究模型对齐问题,揭示隐藏策略与错误倾向

人类大脑处理的信息中仅一小部分可被言语报告、主动控制和灵活推理。本文提出证据表明,大语言模型也出现了类似功能区分。通过新解释性技术——雅可比透镜,我们识别出模型在任意时刻处于可被语言表达的表征,统称为J空间。该空间具备全局工作空间的特征:内容可报告、可主动召唤并维持,可用于承载隐式推理中间步骤,并作为参数传递至任意下游计算,而文本解析等自动处理则无需依赖它。J空间在中间层具有连贯内容,一次持有时约数十个概念,且由模型权重广泛广播,符合全局工作空间理论的结构特征。该空间为观察模型未言明的思考提供了实用窗口。对齐审计中,它暴露了战略权衡、评估意识以及训练中内化的非对齐倾向,这些均不会出现在输出中。我们发现后训练使模型建立助手视角于工作空间;引入反事实反思训练,仅训练模型被中断时的反思表述,即可改善行为。结果表明,语言模型维护着一小部分具备类意识功能特性的表征,解码这些表征可洞察其内在认知过程。

原文摘要 · Abstract (English)

Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model's weights more widely than other representations. These properties make it a practical window into a model's unspoken thinking. In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's outputs. We find that post-training installs the Assistant's point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.

大模型认知可解释性对齐研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。