给大模型加'元认知',让它能自我反思、主动调节推理过程。
Meta-R1: Empowering Large Reasoning Models with Metacognition
- 将推理拆分为任务执行与自我监控两层,实现主动规划与动态调整。
- 在三类基准上最高提升27.3%,同时减少15.7%~32.7%的计算消耗。
- 适用于多种模型和数据集,提升推理稳定性与可迁移性。
大型推理模型(LRMs)在复杂任务中展现出类似人类的思维模式,但其核心缺陷在于缺乏专门的元认知系统——这一人类认知中“思考自身思考”的关键能力。该缺失导致其推理不可控、易出错且缺乏方法论。为此,我们提出Meta-R1,一个基于认知科学原理的通用框架,将推理过程解耦为对象层与元认知层,在级联结构中实现主动规划、在线调控与自适应早停。在三个挑战性基准上的实验表明,Meta-R1在性能上比现有方法最高提升27.3%,令牌消耗降低至原模型的15.7%~32.7%,效率提升达14.8%;同时具备跨数据集与模型骨干的强泛化能力。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) demonstrate remarkable capabilities on complex tasks, exhibiting emergent, human-like thinking patterns. Despite their advances, we identify a fundamental limitation: current LRMs lack a dedicated meta-level cognitive system-an essential faculty in human cognition that enables "thinking about thinking". This absence leaves their emergent abilities uncontrollable (non-adaptive reasoning), unreliable (intermediate error), and inflexible (lack of a clear methodology). To address this gap, we introduce Meta-R1, a systematic and generic framework that endows LRMs with explicit metacognitive capabilities. Drawing on principles from cognitive science, Meta-R1 decomposes the reasoning process into distinct object-level and meta-level components, orchestrating proactive planning, online regulation, and adaptive early stopping within a cascaded framework. Experiments on three challenging benchmarks and against eight competitive baselines demonstrate that Meta-R1 is: (I) high-performing, surpassing state-of-the-art methods by up to 27.3%; (II) token-efficient, reducing token consumption to 15.7% ~ 32.7% and improving efficiency by up to 14.8% when compared to its vanilla counterparts; and (III) transferable, maintaining robust performance across datasets and model backbones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。