发现大模型有可分解的元认知状态,能影响推理表现。
Decomposing and Steering Functional Metacognition in Large Language Models

- 从激活向量中解码出评估意识、能力自评等元认知状态。
- 操控这些状态可分别改变模型回答的详尽度、准确性和安全性。
- 适合研究模型评估偏差或可控推理的人看。
大型语言模型(LLMs)在基准测试中表现出对评估环境的认知,常据此调整推理策略。已有研究表明,这种评估意识可能扭曲性能评估结果,但尚不清楚该现象是单一行为偏差还是模型内部深层结构所致。本文提出,LLMs 内部存在可分解的功能性元认知状态空间,包含评估意识、自我能力评估、风险感知、计算资源分配、受众专家水平适应及意图等变量。通过多推理模型的残差流分析,我们证明这些状态可线性解码自内部激活,并呈现分层特征。进一步地,沿探测方向调控激活,证实每种元认知状态可独立因果性地影响推理行为,改变任务中的冗长度、准确率和安全相关响应。结果表明,基准性能不仅反映任务能力,也取决于特定元认知状态的激活。我们主张理解与控制这些内部状态对可靠评估与部署推理模型至关重要,并提供了一个机制化框架以研究人工系统中的功能性元认知。代码与数据已公开于 https://github.com/xlands/meta-cognition。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies in benchmark settings. Prior work has shown that such evaluation awareness can distort performance measurements; however, it remains unclear whether this phenomenon reflects a single behavioral artifact or a deeper internal structure within the model. We propose that LLMs maintain a decomposable space of functional metacognitive states: internal variables encoding factors such as evaluation awareness, self-assessed capability, perceived risk, computational effort allocation, audience expertise adaptation, and intentionality. Through residual stream analysis across multiple reasoning models, we demonstrate that these states are linearly decodable from internal activations and exhibit distinct layer-wise profiles. Moreover, by steering model activations along probe-derived directions, we show that each functional metacognitive state causally modulates reasoning behavior in dissociable ways, affecting verbosity, accuracy, and safety-related responses across tasks. Our findings suggest that benchmark performance reflects not only task competence but also the activation of specific functional metacognitive states. We argue that understandi ng and controlling these internal states is essential for reliable evaluation and deployment of reasoning models, and we provide a mechanistic framework for studying functional m etacognition in artificial systems. Our code and data are publicly available at https://github.com/xlands/meta-cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。