arXiv:2508.14802cs.AIcs.CL2025-08被引 20

AI能否真正内省?新定义揭示大模型常误判自身状态。

Privileged Self-Access Matters for Introspection in AI

  • 提出更严格的内省定义:需比第三方更可靠地获知内部状态。
  • 实验显示大模型看似内省,实则无法满足新标准。
  • 适合关注AI自我认知能力的开发者与研究者阅读。

AI是否具备内省能力已成为重要现实问题,但目前尚无统一定义。本文基于一种近期提出的‘轻量级’定义,主张采用更严格的定义:若某过程能比同等或更低计算成本的第三方方法更可靠地获取内部状态信息,则可视为内省。通过让大语言模型(LLMs)推理其内部温度参数的实验,我们发现模型虽表现出轻量级内省现象,却无法满足所提严格定义下的真正内省要求。

原文摘要 · Abstract (English)

Whether AI models can introspect is an increasingly important practical question. But there is no consensus on how introspection is to be defined. Beginning from a recently proposed ''lightweight'' definition, we argue instead for a thicker one. According to our proposal, introspection in AI is any process which yields information about internal states through a process more reliable than one with equal or lower computational cost available to a third party. Using experiments where LLMs reason about their internal temperature parameters, we show they can appear to have lightweight introspection while failing to meaningfully introspect per our proposed definition.

AI内省大模型认知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。