arXiv:2503.22233cs.LGcs.AI2025-03被引 8

用熵值自动划分推理步骤,少用标注也能高效训练数学推理模型。

More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty

  • 根据预测熵值自动识别推理关键节点,实现动态分段。
  • 仅用1.5%数据达到顶尖模型效果,推理准确率提升至67.3%。
  • 适合需要降低标注成本的数学推理系统开发人员。

我们提出熵驱动不确定性过程奖励模型(EDU-PRM),一种新型的熵驱动训练框架,用于过程奖励建模,可动态、对齐不确定性的复杂推理步骤分割,无需昂贵的手动步骤标注。与依赖静态划分和人工标注的以往过程奖励模型不同,EDU-PRM在预测熵高的标记处自动锚定步骤边界,有效捕捉内在逻辑转换,并促进对多样化推理路径的高效探索。在ProcessBench基准上,EDU-PRM优于强基线如Math-Shepherd PRM和Omega PRM,且仅使用1.5%训练数据即达到与顶尖模型相当的效果。此外,通过提出的EDU采样策略,生成式推理任务准确率从64.7%提升至67.3%,同时减少32%的令牌使用量。这些结果表明,EDU-PRM是一种可扩展、标注高效的数学推理过程监督范式,为更高效、鲁棒的复杂数学问题求解方法铺平道路。

原文摘要 · Abstract (English)

We introduce the Entropy-Driven Uncertainty Process Reward Model (EDU-PRM), a novel entropy-driven training framework for process reward modeling that enables dynamic, uncertainty-aligned segmentation of complex reasoning steps, eliminating the need for costly manual step annotations. Unlike previous Process Reward Models (PRMs) that rely on static partitioning and human labeling, EDU-PRM automatically anchors step boundaries at tokens with high predictive entropy, effectively capturing intrinsic logical transitions and facilitating efficient exploration of diverse reasoning paths. On the ProcessBench benchmark, EDU-PRM outperforms strong public PRM baselines, such as Math-Shepherd PRM and Omega PRM, and EDU-PRM achieves comparable results with SOTA models while only using 1.5% training data. Furthermore, by leveraging our proposed EDU sampling strategy, we observe accuracy boosts from 64.7% to 67.3% for generative reasoning tasks, accompanied by a reduction of 32% in token usage. These findings underscore the potential of EDU-PRM as a scalable and annotation-efficient paradigm for process supervision in mathematical reasoning, paving the way for more efficient and robust approaches to complex mathematical problem solving.

推理建模奖励模型数学推理无监督分段

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。