arXiv:2602.04145cs.LGcs.CL2026-02被引 1

用少量数据训练出更高效的多模态推理奖励模型

Training Data Efficiency in Multimodal Process Reward Models

  • 基于采样轨迹的标签混合与可靠性设计高效数据筛选方法
  • 仅用10%数据即达全量数据性能,优于随机采样4.1%
  • 适合追求低成本高效训练的多模态模型研究者

多模态过程奖励模型(MPRMs)是实现视觉推理中步骤级监督的核心。现有训练通常依赖大规模蒙特卡洛(MC)标注语料,成本高昂。本文研究MPRM训练的数据效率,发现随机子采样下性能迅速饱和,表明现有MC标注语料存在大量冗余。我们提出理论框架,揭示有效梯度更新依赖于正负步骤的标签混合度与标签可靠性(正向步骤平均MC得分)。据此设计无额外成本的平衡信息评分(BIS),在滚动层面上优先选择兼具混合度与可靠性的样本。在VisualProcessBench上,使用InternVL2.5-8B和Qwen2.5-VL-7B两个主干网络,BIS选取的子集始终匹配甚至超越全量数据表现;尤其仅需10%训练数据即可达到全量性能,相较随机采样相对提升4.1%。

原文摘要 · Abstract (English)

Multimodal Process Reward Models (MPRMs) are central to step-level supervision for visual reasoning in MLLMs. Training MPRMs typically requires large-scale Monte Carlo (MC)-annotated corpora, incurring substantial training cost. This paper studies the data efficiency for MPRM training. Our preliminary experiments reveal that MPRM training quickly saturates under random subsampling of the training data, indicating substantial redundancy within existing MC-annotated corpora. To explain this, we formalize a theoretical framework and reveal that informative gradient updates depend on two factors: label mixtures of positive/negative steps and label reliability (average MC scores of positive steps). Guided by these insights, we propose the Balanced-Information Score (BIS), which prioritizes both mixture and reliability based on existing MC signals at the rollout level, without incurring any additional cost. Across two backbones (InternVL2.5-8B and Qwen2.5-VL-7B) on VisualProcessBench, BIS-selected subsets consistently match and even surpass the full-data performance at small fractions. Notably, the BIS subset reaches full-data performance using only 10% of the training data, improving over random subsampling by a relative 4.1%.

多模态奖励模型数据效率采样策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。