arXiv:2505.07558cs.LGcs.CL2025-05ICML被引 7

不依赖偏好假设,直接优化语言模型输出分布的对齐方法。

Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models

  • 直接估计优选与非优选输出的密度比,避免建模人类偏好。
  • 理论证明随数据量增大,能收敛到真实偏好分布。
  • 适合追求高可靠性对齐的科研与工业应用。

将大语言模型(LLMs)对齐人类偏好对于安全部署至关重要,但现有方法依赖于特定偏好模型(如Bradley-Terry模型),导致统计不一致——数据越多未必更接近真实偏好。为解决这一关键问题,我们提出一种新对齐方法:直接密度比优化(DDRO)。DDRO直接估计优选与非优选输出分布之间的密度比,无需显式建模人类偏好。理论上证明,DDRO具有统计一致性,随着数据规模增加,可收敛至真实偏好分布,且不依赖底层偏好结构。实验表明,DDRO在多个主流基准上表现优于现有方法。该方法实现了真正数据驱动的对齐,为构建更可靠、更符合人类意图的LLM铺平道路。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) with human preferences is crucial for safe deployment, yet existing methods assume specific preference models like Bradley-Terry model. This assumption leads to statistical inconsistency, where more data doesn't guarantee convergence to true human preferences. To address this critical gap, we introduce a novel alignment method Direct Density Ratio Optimization (DDRO). DDRO directly estimates the density ratio between preferred and unpreferred output distributions, circumventing the need for explicit human preference modeling. We theoretically prove that DDRO is statistically consistent, ensuring convergence to the true preferred distribution as the data size grows, regardless of the underlying preference structure. Experiments demonstrate that DDRO achieves superior performance compared to existing methods on many major benchmarks. DDRO unlocks the potential for truly data-driven alignment, paving the way for more reliable and human-aligned LLMs.

模型对齐密度比统计一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。