用双重共识筛选噪声数据,训练出可靠的生物推理过程评分模型。
DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
- 通过自一致与邻域一致双重机制划分数据可信度等级
- 在无需专家标注的情况下实现复杂推理任务的稳定训练
- 适合需要高质量推理评估但标注成本高的研究场景
在科学推理任务中,推理过程的真实性与最终结果同样重要。尽管过程奖励模型(PRMs)可缓解结果奖励模型(ORMs)的粗粒度监督问题,但其应用受限于获取专家验证的逐步标签所带来的高昂成本。本文提出一种基于丰富但噪声较大的弱监督数据训练可靠PRMs的方法。现有弱到强泛化理论缺乏从噪声数据中选取高质量训练信号的具体指导。为此,我们提出双共识弱到强(DC-W2S)框架:通过交叉弱监督者间的自一致(SC)指标与嵌入空间中的邻域一致(NC)指标,将监督信号分层为不同可靠性区间;并采用实例级平衡采样与标签级可靠性感知掩码的课程学习策略引导训练。实验表明,DC-W2S可在无需大量专家标注的前提下训练出稳健的PRM,证明战略性数据筛选比对大规模噪声数据的盲目训练更有效。
原文摘要 · Abstract (English)
In scientific reasoning tasks, the veracity of the reasoning process is as critical as the final outcome. While Process Reward Models (PRMs) offer a solution to the coarse-grained supervision problems inherent in Outcome Reward Models (ORMs), their deployment is hindered by the prohibitive cost of obtaining expert-verified step-wise labels. This paper addresses the challenge of training reliable PRMs using abundant but noisy "weak" supervision. We argue that existing Weak-to-Strong Generalization (W2SG) theories lack prescriptive guidelines for selecting high-quality training signals from noisy data. To bridge this gap, we introduce the Dual-Consensus Weak-to-Strong (DC-W2S) framework. By intersecting Self-Consensus (SC) metrics among weak supervisors with Neighborhood-Consensus (NC) metrics in the embedding space, we stratify supervision signals into distinct reliability regimes. We then employ a curriculum of instance-level balanced sampling and label-level reliability-aware masking to guide the training process. We demonstrate that DC-W2S enables the training of robust PRMs for complex reasoning without exhaustive expert annotation, proving that strategic data curation is more effective than indiscriminate training on large-scale noisy datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。