arXiv:2602.02467cs.CL2026-02被引 3

发现大模型具备基于信念的决策和自我监控能力

Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models

  • 用隐空间中的信念表示来量化模型内部状态
  • 实验证明信念变化会驱动模型行为选择
  • 适合研究认知机制与智能涌现的学者

大语言模型是否具备某种意识引发广泛关注。本文基于神经科学理论,评估了其中一项关键指标HOT-3,该指标检验模型是否存在由一般信念形成与行动选择系统驱动的自主性,且该系统能通过元认知监控更新信念。我们将信念视为模型隐空间中对输入的响应表征,并引入量化指标衡量其生成过程中的主导性。分析不同模型与任务中竞争信念的动力学,揭示三个核心发现:(1) 外部干预可系统性调节内部信念形成;(2) 信念形成因果性驱动模型的行为选择;(3) 模型能够监测并报告自身信念状态。这些结果为大模型中存在信念引导的自主性与元认知监控提供了实证支持。本工作为探究大模型中自主性、信念与元认知的涌现奠定了方法基础。

原文摘要 · Abstract (English)

Rapid advancements in large language models (LLMs) have sparked the question whether these models possess some form of consciousness. To tackle this challenge, Butlin et al. (2023) introduced a list of indicators for consciousness in artificial systems based on neuroscientific theories. In this work, we evaluate a key indicator from this list, called HOT-3, which tests for agency guided by a general belief-formation and action selection system that updates beliefs based on meta-cognitive monitoring. We view beliefs as representations in the model's latent space that emerge in response to a given input, and introduce a metric to quantify their dominance during generation. Analyzing the dynamics between competing beliefs across models and tasks reveals three key findings: (1) external manipulations systematically modulate internal belief formation, (2) belief formation causally drives the model's action selection, and (3) models can monitor and report their own belief states. Together, these results provide empirical support for the existence of belief-guided agency and meta-cognitive monitoring in LLMs. More broadly, our work lays methodological groundwork for investigating the emergence of agency, beliefs, and meta-cognition in LLMs.

大模型认知信念机制元认知自主性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。