arXiv:2606.32032cs.CLcs.AI2026-06被引 1

用元认知反馈强化学习,让大模型更真实地表达不确定性。

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

论文配图:Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
图 1 · 摘自论文原文
  • 通过自评表现优化排序,提升模型自我判断的准确性
  • 在多个任务上实现领先级的不确定性校准,准确率不下降
  • 适合需要可信预测、高可靠性场景的AI系统开发者

元认知是智能的核心能力,指对自身认知过程的监控与调节。然而大模型在关键元认知能力上存在系统性缺陷:高自信下幻觉频发、无法识别知识边界、错误表达内部不确定性,严重削弱可信度。本文提出基于元认知反馈的强化学习(RLMF),通过模型自评表现来优化生成结果的排名,并利用自评结果筛选高质量训练样本,优于传统主动学习。针对‘真实校准’(FC)这一本质元认知任务——使模型表达的不确定性与其内在置信度一致——采用两阶段方法:先校准自报置信度的忠实性,再通过针对性输出编辑映射为自然语言中的情境适配性不确定表达。大量实验表明,RLMF在多种任务上实现通用且顶尖的FC性能,同时保持原有准确率;相比标准强化学习提升最高达63%,显著增强模型对自身能力边界的评估与表达能力。这表明元认知表现可作为有效的强化学习信号,突破以往内在反馈方法的局限,为提升大模型的元认知能力与对齐性提供新范式。

原文摘要 · Abstract (English)

Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes. Yet LLMs exhibit systemic deficiencies in key metacognitive faculties: they hallucinate with high confidence, fail to recognize knowledge boundaries, and misrepresent their internal uncertainty--undermining trustworthiness and reliability. Since monitoring task performance and adapting behavior accordingly are central to metacognition, we posit that models capable of accurately judging their own performance are better positioned to improve it. We operationalize this idea via two novel mechanisms: reinforcement learning with metacognitive feedback (RLMF), a paradigm to refine completion rankings during preference optimization based on the quality of a model's self-judgments of performance, and metacognitive data selection, which uses similar self-judgments to identify high-value training examples, outperforming naive active learning. We apply these innovations to the problem of faithful calibration (FC), a task that is itself fundamentally metacognitive: the goal is to align expressed with intrinsic uncertainty, difficult even for frontier LLMs. We adopt a two-stage, decoupled approach, first using these methods to calibrate the faithfulness of models' self-reported confidence scores, then mapping to natural, context-adaptable linguistic uncertainty via targeted output editing. Extensive experiments show RLMF achieves generalizable, state-of-the-art FC on diverse tasks while preserving accuracy. Further, RLMF surpasses standard RL by up to 63% while enhancing models' ability to assess and express their own capability limits. This positions RLMF as a promising paradigm to enhance LLM metacognition toward improved abilities and alignment, and suggests metacognitive performance as an effective RL signal to overcome limits of prior intrinsic feedback methods.

元认知不确定性强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。