用大模型多方面能力提升奖励模型训练效果
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
- 让教师模型发挥改写、打分、生成三重作用
- 在多个基准上性能超越传统方法,强化学习效果更好
- 适合研究大模型对齐与奖励建模的学者参考
奖励模型(RMs)在对齐大型语言模型(LLMs)与人类偏好方面起着关键作用。由于高质量人类偏好标注难以获取,从生成式大模型中蒸馏偏好已成为标准做法。然而,现有方法大多将教师模型视为简单的二值标注器,未能充分挖掘其在奖励模型蒸馏中的多重潜力。为此,我们提出RM-Distiller框架,系统性地利用教师大模型的三大能力:(1) 优化能力,生成高度相关且对比性强的响应对,提供细粒度信号;(2) 打分能力,通过感知间隔的优化目标引导奖励模型捕捉精确的偏好强度;(3) 生成能力,引入教师模型的生成分布,正则化奖励模型以保留基本语言知识。大量实验表明,RM-Distiller在奖励模型基准和基于强化学习的对齐任务中均显著优于传统蒸馏方法,证明充分利用教师模型的多面能力对有效奖励建模至关重要。据我们所知,这是首个针对生成式大模型进行奖励模型蒸馏的系统性研究。
原文摘要 · Abstract (English)
Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. Due to the difficulty of obtaining high-quality human preference annotations, distilling preferences from generative LLMs has emerged as a standard practice. However, existing approaches predominantly treat teacher models as simple binary annotators, failing to fully exploit the rich knowledge and capabilities for RM distillation. To address this, we propose RM-Distiller, a framework designed to systematically exploit the multifaceted capabilities of teacher LLMs: (1) Refinement capability, which synthesizes highly correlated response pairs to create fine-grained and contrastive signals. (2) Scoring capability, which guides the RM in capturing precise preference strength via a margin-aware optimization objective. (3) Generation capability, which incorporates the teacher's generative distribution to regularize the RM to preserve its fundamental linguistic knowledge. Extensive experiments demonstrate that RM-Distiller significantly outperforms traditional distillation methods both on RM benchmarks and reinforcement learning-based alignment, proving that exploiting multifaceted teacher capabilities is critical for effective reward modeling. To the best of our knowledge, this is the first systematic research on RM distillation from generative LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。