针对奖励模型数据稀缺问题,提出自适应增强方法提升对齐效果。
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
- 根据偏好样本的置信度差异动态分配增强资源
- 在三个数据集上平均提升RewardBench得分与对齐胜率
- 适合低资源场景下需要可靠奖励模型的研究者
奖励建模是强化学习人类反馈(RLHF)、基于RLAIF和PPO对齐的核心,但其可靠性常受限于稀少且异构的人类偏好数据。本文提出MARS(Margin and Semantic-Aware Data Augmentation for Reward Modeling),一种面向低资源奖励建模的自适应数据增强框架。MARS将更多增强资源分配给低置信度偏好对,并通过语义距离引导的精炼机制,在生成合成偏好样本前增强优选-次选对比。在三个偏好数据集、两种奖励模型主干网络及下游对齐评估中,MARS在平均RewardBench表现和对齐胜率上均优于均匀增强、WoN及AdaBoost风格基线。消融实验与独立评委评估表明,性能提升并非仅由语义精炼或GPT-4.1评判耦合造成。
原文摘要 · Abstract (English)
Reward modeling is central to RLHF, RLAIF, and PPO-based alignment, but its reliability is often limited by scarce and heterogeneous human preference data. In this paper, we introduce MARS (Margin and Semantic-Aware Data Augmentation for Reward Modeling), an adaptive augmentation framework for controlled low-resource reward modeling. MARS allocates more augmentation to low-margin preference pairs and uses semantic-distance-based refinement to improve chosen-rejected contrast before generating synthetic preference samples. Across three preference datasets, two reward-model backbones, and downstream alignment evaluations, MARS improves average RewardBench performance and alignment win rates over uniform augmentation, WoN, and AdaBoost-style baselines. Ablations and independent-judge evaluations suggest that the gains are not solely explained by semantic refinement alone or GPT-4.1 judge coupling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。