arXiv:2601.22648cs.AIcs.LG2026-01中稿 · ICML

让大模型学会识别不确定,避免盲目自信出错。

UCPO: Uncertainty-Aware Policy Optimization

  • 用三元优势解耦分离确定与不确定的决策路径,消除奖励偏差。
  • 动态调整不确定性权重,随模型进化和任务难度实时优化。
  • 在数学推理等任务中显著提升模型可靠性,突破知识边界。

构建可信大语言模型的关键在于赋予其内在的不确定性表达能力,从而缓解高风险应用中的过度自信问题。然而,现有基于强化学习的范式(如GRPO)常因二元决策空间和静态不确定性奖励,导致优势偏差,引发过度保守或过度自信。本文揭示了当前引入不确定性奖励的强化学习范式中奖励欺骗与过度自信的根本原因,并提出不确定性感知策略优化(UCPO)框架。UCPO采用三元优势解耦机制,将确定性与不确定性轨迹分离并独立归一化,有效消除优势偏差;同时引入动态不确定性奖励调节机制,根据模型演进和实例难度实时调整不确定性权重。实验结果表明,UCPO在数学推理与通用任务中均能有效缓解奖励失衡,显著提升模型在知识边界外的可靠性。

原文摘要 · Abstract (English)

The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitigating overconfident errors in high-stakes applications. However, existing RL paradigms such as GRPO often suffer from Advantage Bias due to binary decision spaces and static uncertainty rewards, inducing either excessive conservatism or overconfidence. To tackle this challenge, this paper unveils the root causes of reward hacking and overconfidence in current RL paradigms incorporating uncertainty-based rewards, based on which we propose the UnCertainty-Aware Policy Optimization (UCPO) framework. UCPO employs Ternary Advantage Decoupling to separate and independently normalize deterministic and uncertain rollouts, thereby eliminating advantage bias. Furthermore, a Dynamic Uncertainty Reward Adjustment mechanism adapts uncertainty weights in real-time according to model evolution and instance difficulty. Experimental results in mathematical reasoning and general tasks demonstrate that UCPO effectively resolves the reward imbalance, significantly improving the reliability of the model beyond their knowledge boundaries.

强化学习大模型不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。