arXiv:2503.11701cs.LG2025-03综述被引 34

DPO无需奖励模型,直接优化大模型对齐人类偏好。

A Survey of Direct Preference Optimization

  • 用人类偏好直接优化模型,跳过复杂奖励建模
  • 在多个基准上验证了高效稳定且性能接近传统方法
  • 适合想快速部署对齐系统的研究人员和工程师

大型语言模型(LLMs)展现出前所未有的生成能力,但其与人类价值观的对齐仍是确保有益无害部署的关键。尽管基于人类反馈的强化学习(RLHF)已成为对齐LLMs的重要范式,但其依赖复杂的奖励建模,带来了计算效率与训练稳定性之间的权衡。在此背景下,直接偏好优化(DPO)作为简化替代方案迅速兴起,它直接利用人类偏好优化LLMs,无需显式奖励建模。凭借理论简洁性和计算高效性,DPO吸引了大量研究探索其多种实现与应用。然而,该领域尚缺乏系统梳理与对比分析。本文对DPO进行全面综述,提出新颖分类体系,将已有工作分为数据策略、学习框架、约束机制和模型属性四个维度,并在标准化基准上开展严谨的实证分析。此外,还讨论了实际应用场景、开放挑战与未来方向。本工作为理解DPO提供概念框架,也为实践者提供指导,旨在推动更鲁棒、通用的对齐范式发展。所有资源已收集并持续更新于 https://github.com/liushunyu/awesome-direct-preference-optimization。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated unprecedented generative capabilities, yet their alignment with human values remains critical for ensuring helpful and harmless deployments. While Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful paradigm for aligning LLMs with human preferences, its reliance on complex reward modeling introduces inherent trade-offs in computational efficiency and training stability. In this context, Direct Preference Optimization (DPO) has recently gained prominence as a streamlined alternative that directly optimizes LLMs using human preferences, thereby circumventing the need for explicit reward modeling. Owing to its theoretical elegance and computational efficiency, DPO has rapidly attracted substantial research efforts exploring its various implementations and applications. However, this field currently lacks systematic organization and comparative analysis. In this survey, we conduct a comprehensive overview of DPO and introduce a novel taxonomy, categorizing previous works into four key dimensions: data strategy, learning framework, constraint mechanism, and model property. We further present a rigorous empirical analysis of DPO variants across standardized benchmarks. Additionally, we discuss real-world applications, open challenges, and future directions for DPO. This work delivers both a conceptual framework for understanding DPO and practical guidance for practitioners, aiming to advance robust and generalizable alignment paradigms. All collected resources are available and will be continuously updated at https://github.com/liushunyu/awesome-direct-preference-optimization.

大模型对齐直接偏好优化强化学习技术综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。