发现大模型能监控并调控自身神经激活,但能力有限。
Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
- 用类脑反馈机制,通过上下文示例量化模型自省能力
- 监控能力受示例数、激活方向可解释性及方差影响
- 仅能感知极小部分神经激活,对安全防护有重要意义
大型语言模型(LLMs)有时能报告其解题策略,有时却无法识别自身行为的驱动因素,表明其具备有限的元认知能力——即监控自身认知过程以进行报告和自我调控。元认知有助于模型解决复杂任务,但也带来安全风险,因模型可能隐藏内部过程以规避基于神经激活的安全检测。为深入理解此能力,我们引入一种受神经科学启发的神经反馈范式,利用上下文学习量化LLMs报告与控制激活模式的元认知能力。结果表明,该能力受多个因素影响:上下文示例数量、待报告/控制的神经激活方向的语义可解释性,以及该方向解释的方差。这些激活方向构成一个维度远低于模型神经空间的“元认知空间”,说明模型仅能监控其神经激活中的一小部分。该范式为量化LLM元认知提供了实证依据,对人工智能安全(如对抗攻击与防御)具有重要意义。
原文摘要 · Abstract (English)
Large language models (LLMs) can sometimes report the strategies they actually use to solve tasks, yet at other times seem unable to recognize those strategies that govern their behavior. This suggests a limited degree of metacognition - the capacity to monitor one's own cognitive processes for subsequent reporting and self-control. Metacognition enhances LLMs' capabilities in solving complex tasks but also raises safety concerns, as models may obfuscate their internal processes to evade neural-activation-based oversight (e.g., safety detector). Given society's increased reliance on these models, it is critical that we understand their metacognitive abilities. To address this, we introduce a neuroscience-inspired neurofeedback paradigm that uses in-context learning to quantify metacognitive abilities of LLMs to report and control their activation patterns. We demonstrate that their abilities depend on several factors: the number of in-context examples provided, the semantic interpretability of the neural activation direction (to be reported/controlled), and the variance explained by that direction. These directions span a "metacognitive space" with dimensionality much lower than the model's neural space, suggesting LLMs can monitor only a small subset of their neural activations. Our paradigm provides empirical evidence to quantify metacognition in LLMs, with significant implications for AI safety (e.g., adversarial attack and defense).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。