用自然反馈生成连续偏好轨迹,提升大模型对齐效果
ARF-RLHF: Adaptive Reward-Following for RLHF through Emotion-Driven Self-Supervision and Trace-Biased Dynamic Optimization
- 将用户自由反馈转为连续偏好信号,避免二值标签限制
- 在多种大模型和领域上比PPO/DPO提升最多7.6%对齐效果
- 适合追求个性化对齐与可解释性强化学习的研究者
当前基于人类偏好的强化学习(RLHF)方法如PPO和DPO通常将人类偏好简化为二元标签,获取成本高且难以反映个体差异。我们发现,满意与不满意表达在不同用户间呈现稳定的语言模式,表明可从自由文本反馈中提取更丰富的监督信号。基于此,我们提出自适应奖励跟随(ARF),将自然反馈转化为连续偏好轨迹,并通过创新的TraceBias算法进行优化。在多种大模型和偏好领域中,ARF始终优于PPO和DPO,对齐性能最高提升7.6%。结果表明,连续奖励建模为实现个性化且理论严谨的RLHF提供了可扩展路径。
原文摘要 · Abstract (English)
Current RLHF methods such as PPO and DPO typically reduce human preferences to binary labels, which are costly to obtain and too coarse to reflect individual variation. We observe that expressions of satisfaction and dissatisfaction follow stable linguistic patterns across users, indicating that more informative supervisory signals can be extracted from free-form feedback. Building on this insight, we introduce Adaptive Reward-Following (ARF), which converts natural feedback into continuous preference trajectories and optimizes them using the novel TraceBias algorithm. Across diverse LLMs and preference domains, ARF consistently outperforms PPO and DPO, improving alignment by up to 7.6%. Our results demonstrate that continuous reward modeling provides a scalable path toward personalized and theoretically grounded RLHF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。