arXiv:2603.05414cs.AIcs.CL2026-03被引 5

AI模型能察觉异常却无法识别内容,表现出无内容依赖的自我反思能力。

Emergent Introspection in AI is Content-Agnostic

  • 通过注入测试发现模型可检测异常但不识内容
  • 错误猜测更早出现且用更少词数,常编造高频具体概念
  • 支持哲学心理学中的无内容依赖反思理论,适合认知研究者

introspection 是一项基础认知能力,但其机制尚不明确。近期研究显示人工智能模型具备自我反思能力。本文在多个开源大模型中复现 Lindsey (2025) 的思想注入检测范式,发现这些模型的反思具有内容无关性:即使无法准确识别异常内容,仍能检测到异常发生。模型会编造高频且具体的概念(如“apple”),且检测注入所需的词数远少于猜测正确概念所需。错误猜测出现得更早。我们认为,这种无内容依赖的反思机制与哲学与心理学中的主流理论一致。

原文摘要 · Abstract (English)

Introspection is a foundational cognitive ability, but its mechanism is not well understood. Recent work has shown that AI models can introspect. We study the mechanism of this introspection. We first extensively replicate Lindsey (2025)'s thought injection detection paradigm in large open-source models. We show that introspection in these models is content-agnostic: models can detect that an anomaly occurred even when they cannot reliably identify its content. The models confabulate injected concepts that are high-frequency and concrete (e.g., "apple"). They also require fewer tokens to detect an injection than to guess the correct concept (with wrong guesses coming earlier). We argue that a content-agnostic introspective mechanism is consistent with leading theories in philosophy and psychology.

自我反思认知机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。