arXiv:2506.09532cs.LGcs.AI2025-06被引 17

用少量数据训练出精准评估多模态推理步骤的奖励模型。

Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models

  • 通过弱强模型预测一致性筛选可靠推理标签,降低标注成本。
  • 仅用5000样本即在多个基准上达到顶尖性能,提升7.1至10.2分。
  • 适合需要高效推理评估的多模态AI研究与应用开发者。

我们提出Athena-PRM,一种用于评估复杂推理问题中每一步奖励得分的多模态过程奖励模型(PRM)。传统高性能PRM需大量时间与资金投入,主要因推理步骤需逐步标注。现有自动化标注方法如蒙特卡洛估计常产生噪声标签且计算成本高。为此,我们提出利用弱模型与强模型生成结果的一致性作为筛选高质量过程标签的标准。Athena-PRM仅需5000样本即可在多种场景与基准上表现优异。此外,我们还提出两种提升PRM性能的策略:ORM初始化与负样本上采样。我们在三个具体场景中验证:测试时缩放的验证、推理步骤正确性的直接评估、基于奖励排序的微调。Athena-PRM在多个基准和场景中持续领先。当以Qwen2.5-VL-7B为策略模型时,在WeMath上提升10.2分,MathVista上提升7.1分。在VisualProcessBench上达到新的最先进水平,比前序最优结果提升3.9 F1-score,展现其准确评估推理步骤的能力。进一步地,使用Athena-PRM作为奖励模型,通过奖励排序微调构建Athena-7B,在五个基准上显著超越基线。

原文摘要 · Abstract (English)

We present Athena-PRM, a multimodal process reward model (PRM) designed to evaluate the reward score for each step in solving complex reasoning problems. Developing high-performance PRMs typically demands significant time and financial investment, primarily due to the necessity for step-level annotations of reasoning steps. Conventional automated labeling methods, such as Monte Carlo estimation, often produce noisy labels and incur substantial computational costs. To efficiently generate high-quality process-labeled data, we propose leveraging prediction consistency between weak and strong completers as a criterion for identifying reliable process labels. Remarkably, Athena-PRM demonstrates outstanding effectiveness across various scenarios and benchmarks with just 5,000 samples. Furthermore, we also develop two effective strategies to improve the performance of PRMs: ORM initialization and up-sampling for negative data. We validate our approach in three specific scenarios: verification for test time scaling, direct evaluation of reasoning step correctness, and reward ranked fine-tuning. Our Athena-PRM consistently achieves superior performance across multiple benchmarks and scenarios. Notably, when using Qwen2.5-VL-7B as the policy model, Athena-PRM enhances performance by 10.2 points on WeMath and 7.1 points on MathVista for test time scaling. Furthermore, Athena-PRM sets the state-of-the-art (SoTA) results in VisualProcessBench and outperforms the previous SoTA by 3.9 F1-score, showcasing its robust capability to accurately assess the correctness of the reasoning step. Additionally, utilizing Athena-PRM as the reward model, we develop Athena-7B with reward ranked fine-tuning and outperforms baseline with a significant margin on five benchmarks.

多模态推理奖励模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。