arXiv:2502.11520cs.CL2025-02KDD被引 20

AURORA自动化训练通用推理评分模型,提升复杂思维过程评估准确率

AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification

  • 采用集成提示与逆向验证双阶段框架自动标注推理过程
  • 在多样策略分布和长思维链下,评分准确率显著提升
  • 适合需要高精度推理评估的AI系统研发人员使用

先进大语言模型(如o1)的推理能力已革新人工智能应用。然而,由于策略分布多样且人工评估存在局限,复杂推理过程的评估与优化仍是重大挑战。本文提出AURORA,一种基于集成提示与逆向验证的自动化通用过程奖励模型(PRM)训练框架。该框架采用两阶段设计:第一阶段通过多样化提示策略与集成方法,实现推理过程的自动化标注与评估,确保奖励学习的稳健性;第二阶段利用实际参考答案进行逆向验证,增强模型输出验证能力,提升训练精度。为评估性能,我们扩展现有ProcessBench基准,引入UniversalBench,该基准在多种策略分布下,对包含长思维链(CoT)输出的完整推理轨迹进行评测。实验表明,AURORA显著提升过程评估准确率,改善了不同策略分布及长CoT响应下的PRM表现。项目将开源,Universal-PRM-7B模型可在HuggingFace获取。

原文摘要 · Abstract (English)

The reasoning capabilities of advanced large language models (LLMs) like o1 have revolutionized artificial intelligence applications. Nevertheless, evaluating and optimizing complex reasoning processes remain significant challenges due to diverse policy distributions and the inherent limitations of human effort and accuracy. In this paper, we present AURORA, a novel automated framework for training universal process reward models (PRMs) using ensemble prompting and reverse verification. The framework employs a two-phase approach: First, it uses diverse prompting strategies and ensemble methods to perform automated annotation and evaluation of processes, ensuring robust assessments for reward learning. Second, it leverages practical reference answers for reverse verification, enhancing the model's ability to validate outputs and improving training accuracy. To assess the framework's performance, we extend beyond the existing ProcessBench benchmark by introducing UniversalBench, which evaluates reward predictions across full trajectories under diverse policy distribtion with long Chain-of-Thought (CoT) outputs. Experimental results demonstrate that AURORA enhances process evaluation accuracy, improves PRMs' accuracy for diverse policy distributions and long-CoT responses. The project will be open-sourced at https://auroraprm.github.io/. The Universal-PRM-7B is available at https://huggingface.co/infly/Universal-PRM-7B.

推理评估奖励模型自动化训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。