解析强化学习与直接偏好优化的性能差异,揭示适用场景。
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
- 从理论拆解性能差距:显式表示与隐式表示两类原因。
- 在样本有限时,两阶段方法比直接优化更少需要数据。
- 当模型均错误设定时,在线直接偏好优化表现最佳。
本文对两阶段人类反馈强化学习(RLHF)与直接偏好优化(DPO)之间的性能差距进行了细粒度的理论分析。研究将该差距分解为两种来源:精确优化下的显式表示差距,以及有限样本下的隐式表示差距。在精确优化条件下,我们刻画了奖励模型与策略模型类相对容量如何影响最终策略质量,表明在不同模型误设情况下,RLHF、DPO 或在线 DPO 可能各有优劣。值得注意的是,当奖励与策略模型类同构且均误设时,在线 DPO 表现优于标准 RLHF 和 DPO。在近似优化设置中,我们构造了一个真实奖励稀疏的例子,证明 RLHF 恢复有效奖励模型所需样本显著少于 DPO,凸显两阶段学习的统计优势。这些结果全面揭示了两种方法在不同设定下的性能差异,并为实际选择提供了指导。
原文摘要 · Abstract (English)
We present a fine-grained theoretical analysis of the performance gap between two-stage reinforcement learning from human feedback~(RLHF) and direct preference optimization~(DPO). Our study decomposes this gap into two sources: the explicit representation gap under exact optimization and the implicit representation gap under finite samples. In the exact optimization setting, we characterize how the relative capacities of the reward and policy model classes influence the final policy qualities. We show that RLHF, DPO, or online DPO can outperform one another depending on type of model mis-specifications. Notably, online DPO can outperform both RLHF and standard DPO when the reward and policy model classes are isomorphic and both mis-specified. In the approximate optimization setting, we provide a concrete construction where the ground-truth reward is sparse and show that RLHF requires significantly fewer samples than DPO to recover an effective reward model, highlighting a statistical advantage of two-stage learning. Together, these results provide a comprehensive understanding of the performance gap between RLHF and DPO under various settings, and offer practical insights into when each method is preferred.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。