评测大模型在动态环境中的认知自主能力,发现其反思能力严重不足。
Reflection-Bench: Evaluating Epistemic Agency in Large Language Models
- 基于认知心理学设计七维任务,评估模型自主构建与更新信念的能力。
- 16个模型分三层次表现,顶尖模型仍缺乏深度自我反思能力。
- 适合关注AI Agent认知可靠性与模型内省机制的研究者。
随着大型语言模型(LLMs)越来越多地被用作人工智能代理的认知引擎,其可靠性与有效性关键取决于内在的认知自主性,而这一能力尚未得到充分研究。认知自主性指模型在动态环境中灵活构建、调整和监控自身信念的能力,是一种不依赖特定工具或应用的基础模型能力。我们系统刻画了认知自主性的七个相互关联维度:预测、决策、感知、记忆、反事实思维、信念更新与元反思。据此提出Reflection-Bench,一个受认知心理学启发的基准测试,包含七项具有长期相关性且最小化数据泄露的任务。通过对16个模型使用三种提示策略进行综合评估,识别出明显的三层性能层级,并揭示当前大模型在元反思方面存在显著局限。尽管顶尖模型展现出初步的认知自主迹象,但研究结果提示若干未来方向:增强核心认知功能、改善跨功能协同、发展自适应处理机制。代码与数据已公开于https://github.com/AI45Lab/ReflectionBench。
原文摘要 · Abstract (English)
With large language models (LLMs) increasingly deployed as cognitive engines for AI agents, the reliability and effectiveness critically hinge on their intrinsic epistemic agency, which remains understudied. Epistemic agency, the ability to flexibly construct, adapt, and monitor beliefs about dynamic environments, represents a base-model-level capacity independent of specific tools, modules, or applications. We characterize the holistic process underlying epistemic agency, which unfolds in seven interrelated dimensions: prediction, decision-making, perception, memory, counterfactual thinking, belief updating, and meta-reflection. Correspondingly, we propose Reflection-Bench, a cognitive-psychology-inspired benchmark consisting of seven tasks with long-term relevance and minimization of data leakage. Through a comprehensive evaluation of 16 models using three prompting strategies, we identify a clear three-tier performance hierarchy and significant limitations of current LLMs, particularly in meta-reflection capabilities. While state-of-the-art LLMs demonstrate rudimentary signs of epistemic agency, our findings suggest several promising research directions, including enhancing core cognitive functions, improving cross-functional coordination, and developing adaptive processing mechanisms. Our code and data are available at https://github.com/AI45Lab/ReflectionBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。