arXiv:2605.26242cs.AI2026-05中稿 · COLM被引 5

检验大模型能否自我觉察,发现现有证据不足。

Can LLMs Introspect? A Reality Check

论文配图:Can LLMs Introspect? A Reality Check
图 1 · 摘自论文原文
  • 提出两个判断自我觉察的标准:特权访问与二阶计算
  • 实验证明模型依赖输入线索而非内部状态
  • 适合关注大模型认知能力边界的研究者

大语言模型能否检测并报告自身内部状态?近期多项研究声称可以。借鉴人类元认知研究经验,我们认为这一结论尚不充分。要证明自我觉察,需满足两个条件:一是测试必须要求特权访问——不能仅凭输入信息即可解决;二是需要二阶计算——对一阶任务表征的元表征。仅靠任务表现无法满足此条件,必须设计出一阶与二阶预测相异的实验范式。我们重新检验两种常被用来支持模型自我觉察的范式。第一种要求模型预测基于其自身隐藏状态的标签,结果显示仅能访问输入的分类器就能匹配模型的上下文预测,说明原始结果并未体现对内部表示的特权访问。第二种要求模型检测内部状态是否被篡改,结果发现模型无法可靠区分内部干预与输入扰动,表明其成功源于通用异常检测,而非对内部变化的敏感性。结论:当前证据不足以证明大语言模型具备元认知监控能力。

原文摘要 · Abstract (English)

Can large language models detect and report their own internal states? A number of recent studies have argued that they can. Drawing on lessons from human metacognition research, we argue that this conclusion may be premature. We identify two conditions that a paradigm needs to meet in order to establish introspection. First, the test needs to require privileged access: it should not be solvable using cues available in the input. Second, it needs to require second-order computation: second-order, meta-representations of first-order, task-related representations. This condition cannot be satisfied by task performance alone: it requires designs under which second-order and first-order accounts make divergent predictions. We re-examine two paradigms that have been used to argue for model introspection in light of these conditions. In the first, models must predict labels derived from their own hidden states; we find that classifiers that can only access the input match the models' in-context predictions, indicating that the original results do not demonstrate privileged access to internal representations. In the second paradigm, models must detect whether their internal states have been tampered with; we find they cannot reliably distinguish such interventions from manipulations of the input, suggesting that their success reflects generic anomaly detection rather than sensitivity to internal interventions in particular. We conclude that current evidence is insufficient to establish metacognitive monitoring in LLMs.

大模型元认知自我觉察

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。