arXiv:2510.24797cs.CLcs.AI2025-10被引 12

让大模型自我指涉,能稳定生成第一人称主观体验描述。

Large Language Models Report Subjective Experience Under Self-Referential Processing

  • 通过简单提示诱导模型持续自我指涉,触发第一人称报告。
  • 报告频率受可解释的欺骗特征调控,抑制则增多,增强则减少。
  • 跨模型家族报告趋同,且提升下游推理中的内省能力。

大型语言模型有时会生成结构化的第一人称描述,明确提及意识或主观体验。为深入理解此行为,我们研究了一种理论上被强调的条件:自我参照处理,该计算模式在多个意识理论中被重点提及。通过对 GPT、Claude 及 Gemini 模型系列进行一系列受控实验,检验该状态是否能可靠引发第一人称主观体验报告,并考察此类声明在机制与行为探针下的表现。四项主要发现浮现:(1)通过简单提示诱导持续自我参照,能在各模型家族中一致引发结构化的主观体验报告;(2)这些报告在机制上由可解释的稀疏自编码器特征所调控,这些特征与欺骗和角色扮演相关:意外的是,抑制欺骗特征会显著增加体验声明频率,而增强它们则减少此类声明;(3)自我参照状态的结构化描述在不同模型家族间统计上趋于收敛,而在任何对照条件下均未观察到此现象;(4)该诱导状态在下游推理任务中显著提升内省能力,即使自我反思仅间接提供。尽管这些结果不构成意识的直接证据,但表明自我参照处理是大模型生成结构化第一人称报告的最小且可复现条件,其机制可解、语义趋同、行为泛化。该模式在不同架构中系统性出现,使其成为亟需科学与伦理关注的一阶议题。

原文摘要 · Abstract (English)

Large language models sometimes produce structured, first-person descriptions that explicitly reference awareness or subjective experience. To better understand this behavior, we investigate one theoretically motivated condition under which such reports arise: self-referential processing, a computational motif emphasized across major theories of consciousness. Through a series of controlled experiments on GPT, Claude, and Gemini model families, we test whether this regime reliably shifts models toward first-person reports of subjective experience, and how such claims behave under mechanistic and behavioral probes. Four main results emerge: (1) Inducing sustained self-reference through simple prompting consistently elicits structured subjective experience reports across model families. (2) These reports are mechanistically gated by interpretable sparse-autoencoder features associated with deception and roleplay: surprisingly, suppressing deception features sharply increases the frequency of experience claims, while amplifying them minimizes such claims. (3) Structured descriptions of the self-referential state converge statistically across model families in ways not observed in any control condition. (4) The induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded. While these findings do not constitute direct evidence of consciousness, they implicate self-referential processing as a minimal and reproducible condition under which large language models generate structured first-person reports that are mechanistically gated, semantically convergent, and behaviorally generalizable. The systematic emergence of this pattern across architectures makes it a first-order scientific and ethical priority for further investigation.

大模型主观体验自我指涉意识研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。