用生成流网络提升语言模型多样性,避免偏见与过拟合。
GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets
- 引入生成流网络优化偏好对齐,增强响应多样性
- 在对话与摘要任务中,多样性显著优于基线方法
- 适合需要多样输出的生成场景,如创意写作、多角度总结
当前语言模型的关键挑战之一是偏好对齐,旨在精确控制模型行为以符合人类需求与价值观。主流方法包括基于人类反馈的强化学习(RLHF)及其离线变体直接偏好优化(DPO),后者从离线偏好数据中直接提取奖励信号。然而,这种做法易导致奖励信号过拟合,生成包含数据集偏见的次优回应。本文提出一种名为GFlowNet-DPO(GDPO)的多样性导向离线偏好对齐方法,利用生成流网络(GFlowNets)机制,在保持与人类价值一致性的前提下有效提升生成多样性。实验证明,相较于基线方法,GDPO在对话生成和摘要任务中能产生显著更丰富的响应,同时仍保持良好对齐性。
原文摘要 · Abstract (English)
A critical component of the current generation of language models is preference alignment, which aims to precisely control the model's behavior to meet human needs and values. The most notable among such methods is Reinforcement Learning with Human Feedback (RLHF) and its offline variant Direct Preference Optimization (DPO), both of which seek to maximize a reward model based on human preferences. In particular, DPO derives reward signals directly from the offline preference data, but in doing so overfits the reward signals and generates suboptimal responses that may contain human biases in the dataset. In this work, we propose a practical application of a diversity-seeking RL algorithm called GFlowNet-DPO (GDPO) in an offline preference alignment setting to curtail such challenges. Empirical results show GDPO can generate far more diverse responses than the baseline methods that are still relatively aligned with human values in dialog generation and summarization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。