arXiv:2507.07855cs.LGcs.AI2025-07

从人类选择理论出发,揭示DPO训练算法的深层机制与非凸损失的可行性

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable)

  • 基于人类选择理论重构DPO范式,实现更通用的偏好优化框架
  • 证明非凸损失在DPO中同样有效,突破传统凸性依赖
  • 为偏好学习提供统一理论支撑,适合研究强化学习与人机对齐的学者

规范性理论使我们能从基本原理推导出机器学习算法的关键部分,这在当前对机器学习工作日益严格的审视背景下尤为关键。直接偏好优化(DPO)巧妙地绕过奖励建模,通过与特定的人类选择规范模型建立显式联系而实现。本文将这一联系提升至DPO规范框架的普遍性高度。为此,需重新构建人类选择理论的传统路径以更好地适配强化学习与机器学习。该研究拓展了偏好优化的广泛视角,涵盖当前DPO各类后续工作。同时揭示了令人意外的丰富成果:支持非凸损失函数,任何符合要求的机器学习分析选择均可嵌入任意人类选择模型,并构建了一个足够广阔的规范框架,可保护DPO的各种扩展(如边际调整、长度修正等)。文中还提供了一个远离主流DPO场景的简化实验。

原文摘要 · Abstract (English)

Normative theories allow one to elicit key parts of a ML algorithm from first principles, which is crucial at a time of championed scrutiny for ML work. Direct Preference Optimization (DPO) cleverly bypasses reward modeling by making an explicit link with a specific normative model of human choice. Our paper elevates this connection to the full generality of DPO's normative framework. Getting there requires reworking human choice theory's textbook path for a better RLHF/ML fit. It elevates the connection to a remarkably broad viewpoint on preference optimization, considering the current panorama of DPO follow-ups. It also unveils unexpected riches for ML, chief among which the support for non-convex losses, the fact that any compliant ML analytical choice can be embedded with any human choice model, and a normative framework's umbrella wide enough to safeguard DPO's extensions (margins, length correction, ...). A toy experiment ``far away'' from the DPO crowd is given.

偏好优化人类选择DPO非凸损失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。