用最优传输方法提升大模型奖励模型在噪声偏好数据下的鲁棒性
Optimal Transport for LLM Reward Modeling from Noisy Preference

- 基于最优传输构建联合一致性差异,对齐预测分布与偏好数据
- 引入部分传输机制放松质量守恒,自动排除语义不一致的噪声样本
- 理论证明可优化更紧的清洁风险上界,适合处理真实世界噪声数据
奖励模型是人类反馈强化学习(RLHF)的核心,但真实数据集不可避免地包含噪声偏好。传统训练目标易过拟合这些错误,而现有去噪方法常依赖同质噪声假设,难以捕捉语言偏好复杂性。为此,我们提出SelectiveRM框架,基于最优传输原理。首先设计联合一致性差异,使模型预测分布与偏好数据对齐;其次为克服严格质量守恒导致模型必须拟合异常值的局限,引入通过部分传输实现的质量松弛机制,实现对语义不一致噪声样本的自主剔除。理论上,我们证明SelectiveRM优化了未观测清洁风险的更紧上界。大量实验表明,该方法在多个基准测试中显著优于当前最先进基线。
原文摘要 · Abstract (English)
Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training objectives tend to overfit these errors, while existing denoising approaches often rely on homogeneous noise assumptions that fail to capture the complexity of linguistic preferences. To handle these challenges, we propose SelectiveRM, a framework grounded in optimal transport. We first devise a Joint Consistency Discrepancy to align the distribution of model predictions with preference data. Furthermore, to address the limitation of strict mass conservation which compels the model to fit outliers, we incorporate a Mass Relaxation mechanism via partial transport. This enables the autonomous exclusion of samples with noisy preference that contradict semantic consistency. Theoretically, we demonstrate that SelectiveRM optimizes a tighter upper bound on the unobserved clean risk. Extensive experiments validate that our approach significantly outperforms state-of-the-art baselines across diverse benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。