发现大模型有内在元认知能力,通过新方法可更准确评估其自我纠错能力。
Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens
- 提出AutoMeco框架,自动评估大模型的元认知能力
- MIRA策略提升元认知评估效果,在3个数学数据集上验证
- 无需训练,适配研究模型自我反思机制的研究者
以往研究多关注大语言模型(LLMs)在推理链中识别认知错误的能力,但很少探讨其元认知能力(如对步骤错误的自知)。尽管已有自评估指标如困惑度,可反映答案正确性,但缺乏对推理步骤层面的分析与适应。本文提出AutoMeco框架,用于评测现有元认知评估方法,并设计无需训练的马尔可夫内在奖励调整策略MIRA,以增强评估精度。在三个数学推理数据集和三种大模型上的实验表明,AutoMeco评估结果合理,与Best-of-N验证一致;使用MIRA后,元认知评估能力显著提升。
原文摘要 · Abstract (English)
Previous research has primarily focused on the cognitive error detection capabilities of Large Language Models (LLMs), often prompting them to analyze mistakes in reasoning chains. However, few studies have examined the meta-cognitive abilities of LLMs (e.g., their self-awareness of step errors), which are crucial for their reliability. While studies on LLM self-evaluation present some measures, such as perplexity, which can reflect the answer correctness and be viewed as the lens of meta-cognition, they lack step-level analysis and adaptation. This paper studies the evaluation of LLM meta-cognition using the current lenses and how to improve these lenses. Specifically, we propose AutoMeco, an Automated Meta-cognition Evaluation framework for benchmarking the existing lenses. Furthermore, a training-free Markovian Intrinsic Reward Adjustment strategy, MIRA, is proposed to boost current meta-cognition lenses. Experimental results on three mathematical reasoning datasets and three LLMs show the reasonableness of AutoMeco by comparing it with Best-of-N verification. Moreover, the meta-cognition ability of LLMs can be better evaluated using MIRA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。