arXiv:2502.00814cs.LGcs.CL2025-02被引 6

让大模型听话不光看长短,而是分清人类偏好和长度要求。

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling

  • 用响应条件化的布拉德利-特瑞模型分离长度与语义偏好
  • 在多个数据集上显著减少长度偏见,提升长度指令遵循能力
  • 适合关注模型对齐与可控生成的研究者和工程师

基于人类反馈的强化学习(RLHF)通过可学习的奖励模型建模人类偏好,并利用强化学习算法最大化奖励得分,已在对齐大语言模型方面取得显著成效。然而,这些奖励模型易受各种表面混淆因素的干扰,其中长度偏见尤为突出。尽管长度偏见对偏好建模的影响表明大模型对长度感知具有内在敏感性,但初步实验发现微调后的模型仍难以遵守明确的长度指令。为解决这两个问题,我们提出一种新框架,使奖励模型能显式区分人类语义偏好与响应长度需求。具体地,引入响应条件化的布拉德利-特瑞(Rc-BT)模型,通过在增强数据集上训练,提升其缓解长度偏见与遵循长度指令的能力。此外,我们提出了Rc-RM与Rc-DPO算法,利用Rc-BT模型进行奖励建模和直接策略优化,同时缓解长度偏见并促进长度指令遵循。在多个基础模型与数据集上的广泛实验表明该方法有效且具备良好泛化性。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward model and employing a reinforcement learning algorithm to maximize the reward model's scores. However, these reward models are susceptible to exploitation through various superficial confounding factors, with length bias emerging as a particularly significant concern. Moreover, while the pronounced impact of length bias on preference modeling suggests that LLMs possess an inherent sensitivity to length perception, our preliminary investigations reveal that fine-tuned LLMs consistently struggle to adhere to explicit length instructions. To address these two limitations, we propose a novel framework wherein the reward model explicitly differentiates between human semantic preferences and response length requirements. Specifically, we introduce a $\textbf{R}$esponse-$\textbf{c}$onditioned $\textbf{B}$radley-$\textbf{T}$erry (Rc-BT) model that enhances the model's capability in length bias mitigating and length instruction following, through training on our augmented dataset. Furthermore, we propose the Rc-RM and Rc-DPO algorithm to leverage the Rc-BT model for reward modeling and direct policy optimization (DPO) of LLMs, simultaneously mitigating length bias and promoting adherence to length instructions. Extensive experiments across various foundational models and datasets demonstrate the effectiveness and generalizability of our approach.

模型对齐长度偏见强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。