arXiv:2409.11704cs.CLcs.LG2024-09ACL被引 44

研究发现模型偏好特定格式,会误导对齐效果。

From Lists to Emojis: How Format Bias Affects Model Alignment

  • 分析了列表、表情符号等格式偏见对模型对齐的影响。
  • 仅用不到1%的偏置数据就能显著影响奖励模型。
  • 适合关注模型评估与对齐算法设计的研究者。

本文研究强化学习中人类反馈(RLHF)的格式偏见问题。我们发现,主流偏好模型(包括人类评估者、GPT-4及RewardBench排名靠前的模型)普遍存在对特定格式模式(如列表、链接、粗体、表情符号)的强烈偏好。大型语言模型可利用这些偏见,在AlpacaEval和LMSYS Chatbot Arena等基准上获得更高排名。其中,冗长偏见尤为突出:偏好更长的回复,即使其质量与短回复相当甚至更低。本文扩展了对偏好学习中偏见的研究,涵盖除长度外的多种格式偏见。实验表明,仅需不足1%的偏置数据即可显著引入格式偏差。此外,下游对齐算法(如best-of-n采样与在线迭代DPO)也容易被此类偏见操纵,因改变格式比提升内容质量更易实现。研究强调,在设计对齐算法与评估模型时,必须区分格式与内容。

原文摘要 · Abstract (English)

In this paper, we study format biases in reinforcement learning from human feedback (RLHF). We observe that many widely-used preference models, including human evaluators, GPT-4, and top-ranking models on the RewardBench benchmark, exhibit strong biases towards specific format patterns, such as lists, links, bold text, and emojis. Furthermore, large language models (LLMs) can exploit these biases to achieve higher rankings on popular benchmarks like AlpacaEval and LMSYS Chatbot Arena. One notable example of this is verbosity bias, where current preference models favor longer responses that appear more comprehensive, even when their quality is equal to or lower than shorter, competing responses. However, format biases beyond verbosity remain largely underexplored in the literature. In this work, we extend the study of biases in preference learning beyond the commonly recognized length bias, offering a comprehensive analysis of a wider range of format biases. Additionally, we show that with a small amount of biased data (less than 1%), we can inject significant bias into the reward model. Moreover, these format biases can also be easily exploited by downstream alignment algorithms, such as best-of-n sampling and online iterative DPO, as it is usually easier to manipulate the format than to improve the quality of responses. Our findings emphasize the need to disentangle format and content both for designing alignment algorithms and evaluating models.

模型对齐格式偏见偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。