arXiv:2502.11250cs.CL2025-02被引 21

用不确定性量化提升大模型数学推理步骤验证的可靠性

Uncertainty-Aware Step-wise Verification with Generative Reward Models

  • 提出新方法CoT熵,衡量生成式评分模型在每步推理中的不确定性
  • 引入不确定性估计后,判断模型的验证结果更稳定可靠
  • 适合关注大模型推理可信度与安全性的研究者

复杂多步推理任务(如解数学题)对大语言模型仍具挑战。虽常用结果监督,但通过过程奖励模型(PRM)进行过程监督可提供中间奖励,验证解题步骤的正确性。然而,作为人类判断的代理,PRM存在可靠性问题,易受奖励劫持影响。本文提出利用不确定性量化(UQ)来增强生成式奖励模型在数学推理中逐步验证的可靠性。我们引入一种新方法——CoT熵,其在量化PRM不确定性方面优于现有方法。实验表明,结合不确定性估计能显著提升判断型语言模型(judge-LM)PRM的鲁棒性,实现更可靠的验证。

原文摘要 · Abstract (English)

Complex multi-step reasoning tasks, such as solving mathematical problems, remain challenging for large language models (LLMs). While outcome supervision is commonly used, process supervision via process reward models (PRMs) provides intermediate rewards to verify step-wise correctness in solution traces. However, as proxies for human judgement, PRMs suffer from reliability issues, including susceptibility to reward hacking. In this work, we propose leveraging uncertainty quantification (UQ) to enhance the reliability of step-wise verification with generative reward models for mathematical reasoning tasks. We introduce CoT Entropy, a novel UQ method that outperforms existing approaches in quantifying a PRM's uncertainty in step-wise verification. Our results demonstrate that incorporating uncertainty estimates improves the robustness of judge-LM PRMs, leading to more reliable verification.

推理验证不确定性量化生成式评分大模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。