用互信息统一多种强化学习对齐方法,揭示其内在联系。
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
- 基于互信息构建可调节先验的新损失函数
- 证明SimPO、TDPO等多算法可由该框架推导
- 适合对大模型对齐机制感兴趣的研学者
大语言模型(LLMs)的后对齐对提升其实用性、安全性和与人类意图的一致性至关重要。直接偏好优化(DPO)因其能直接基于人类反馈优化模型,已成为最广泛使用的对齐算法之一。然而,文献中存在大量DPO变体,使研究者难以厘清其相互关系。本文提出一种受互信息启发的统一框架,引入具备灵活先验的新损失函数。通过合理设定这些先验,我们证明了包括SimPO、TDPO、SparsePO在内的多个现有算法均可从该框架中推导得出。这一统一视角提供了更清晰、结构化的理解路径,有助于研究者把握不同DPO变体间的关联。我们旨在简化DPO算法生态,促进社区深入洞察,并为开发更稳健、可解释的对齐技术奠定基础。
原文摘要 · Abstract (English)
Post-alignment of large language models (LLMs) is critical in improving their utility, safety, and alignment with human intentions. Direct preference optimisation (DPO) has become one of the most widely used algorithms for achieving this alignment, given its ability to optimise models based on human feedback directly. However, the vast number of DPO variants in the literature has made it increasingly difficult for researchers to navigate and fully grasp the connections between these approaches. This paper introduces a unifying framework inspired by mutual information, which proposes a new loss function with flexible priors. By carefully specifying these priors, we demonstrate that many existing algorithms, such as SimPO, TDPO, SparsePO, and others, can be derived from our framework. This unification offers a clearer and more structured approach, allowing researchers to understand the relationships between different DPO variants better. We aim to simplify the landscape of DPO algorithms, making it easier for the research community to gain insights and foster further advancements in LLM alignment. Ultimately, we hope our framework can be a foundation for developing more robust and interpretable alignment techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。