arXiv:2509.21545cs.LGcs.AI2025-09被引 15

测试大模型能否识别自身认知状态,发现其元认知能力有限且与人类不同。

Evidence for Limited Metacognition in LLMs

  • 通过行为实验评估模型对自身信心的认知和利用能力
  • 前沿大模型能判断答题正确性并预判答案,但精度有限
  • 适合关注AI自我意识与安全的科研人员阅读

大语言模型是否具备元认知能力(即对自身认知状态的觉察)正引发广泛关注,但相关评估方法仍不成熟。本文提出一种量化评估框架,借鉴非人动物元认知研究,不依赖模型自述,而是检验其能否策略性运用内部状态知识。通过两个实验范式,我们发现2024年后推出的前沿模型展现出一定元认知能力,包括评估自身回答事实与推理问题的可信度,以及预判自身输出并合理利用该信息。结合令牌概率分析,表明存在可能支持元认知的上游内部信号。然而这些能力在分辨率上受限、呈现情境依赖性,且在本质特征上与人类明显不同。模型间还存在相似能力下的显著差异,暗示后训练阶段可能影响元认知发展。

原文摘要 · Abstract (English)

The possibility of LLM self-awareness and even sentience is gaining increasing public attention and has major safety and policy implications, but the science of measuring them is still in a nascent state. Here we introduce a novel methodology for quantitatively evaluating metacognitive abilities in LLMs. Taking inspiration from research on metacognition in nonhuman animals, our approach eschews model self-reports and instead tests to what degree models can strategically deploy knowledge of internal states. Using two experimental paradigms, we demonstrate that frontier LLMs introduced since early 2024 show increasingly strong evidence of certain metacognitive abilities, specifically the ability to assess and utilize their own confidence in their ability to answer factual and reasoning questions correctly and the ability to anticipate what answers they would give and utilize that information appropriately. We buttress these behavioral findings with an analysis of the token probabilities returned by the models, which suggests the presence of an upstream internal signal that could provide the basis for metacognition. We further find that these abilities 1) are limited in resolution, 2) emerge in context-dependent manners, and 3) seem to be qualitatively different from those of humans. We also report intriguing differences across models of similar capabilities, suggesting that LLM post-training may have a role in developing metacognitive abilities.

元认知大模型认知评估AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。