arXiv:2410.18640cs.CL2024-10ICLR被引 24

用弱模型指导强模型对齐,实现更优人类偏好对齐效果

Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model

  • 让强模型学习弱模型对齐前后的分布差异,实现偏好迁移
  • 在Arena-Hard上胜率从39.70提升至49.60,显著超越基线
  • 适合希望低成本提升大模型对齐能力的研究者和工程师

语言模型与人类偏好的对齐已成为研究热点,使模型能更好地满足用户需求。受‘弱到强泛化’启发——即在弱模型生成标签上微调的强模型可持续优于其弱监督者——我们将其扩展至模型对齐领域。本工作发现,弱模型的对齐行为可有效迁移到强模型上,并产生放大效应。为此,提出弱到强偏好优化(WSPO)方法,通过学习弱模型对齐前后的分布差异来实现强模型对齐。实验表明,WSPO表现优异:在Arena-Hard上,Qwen2-7B-Instruct的胜率从39.70提升至49.60;在AlpacaEval 2上取得47.04的长度控制胜率;在MT-bench上得分达7.33。结果表明,利用弱模型引导强模型获得高对齐能力是可行的。

原文摘要 · Abstract (English)

Aligning language models (LMs) with human preferences has become a key area of research, enabling these models to meet diverse user needs better. Inspired by weak-to-strong generalization, where a strong LM fine-tuned on labels generated by a weaker model can consistently outperform its weak supervisor, we extend this idea to model alignment. In this work, we observe that the alignment behavior in weaker models can be effectively transferred to stronger models and even exhibit an amplification effect. Based on this insight, we propose a method called Weak-to-Strong Preference Optimization (WSPO), which achieves strong model alignment by learning the distribution differences before and after the alignment of the weak model. Experiments demonstrate that WSPO delivers outstanding performance, improving the win rate of Qwen2-7B-Instruct on Arena-Hard from 39.70 to 49.60, achieving a remarkable 47.04 length-controlled win rate on AlpacaEval 2, and scoring 7.33 on MT-bench. Our results suggest that using the weak model to elicit a strong model with a high alignment ability is feasible.

偏好对齐模型迁移强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。