将强化学习对齐与模仿学习关联,提出统一新框架提升大模型对齐效果。
On a Connection Between Imitation Learning and RLHF
- 从模仿学习视角重构强化学习对齐,揭示其隐含的模仿机制
- 新框架DIL在多个基准上优于现有方法,性能显著提升
- 为模型对齐提供统一理论视角,适合研究对齐算法的学者
本文从模仿学习(IL)的角度研究大语言模型与偏好数据的对齐问题。建立了强化学习从人类反馈(RLHF)与模仿学习之间的紧密理论联系,发现RLHF本质上是在偏好数据分布上执行模仿学习。基于此联系,提出了一个严谨的框架DIL,直接优化模仿学习目标。DIL为对齐提供了统一的模仿学习视角,可涵盖现有对齐算法作为特例,并自然衍生出新的变体。通过连接IL与RLHF,DIL为理解基于人类反馈的对齐提供了新洞见。大量实验表明,DIL在多个挑战性基准上优于现有方法。
原文摘要 · Abstract (English)
This work studies the alignment of large language models with preference data from an imitation learning perspective. We establish a close theoretical connection between reinforcement learning from human feedback RLHF and imitation learning (IL), revealing that RLHF implicitly performs imitation learning on the preference data distribution. Building on this connection, we propose DIL, a principled framework that directly optimizes the imitation learning objective. DIL provides a unified imitation learning perspective on alignment, encompassing existing alignment algorithms as special cases while naturally introducing new variants. By bridging IL and RLHF, DIL offers new insights into alignment with RLHF. Extensive experiments demonstrate that DIL outperforms existing methods on various challenging benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。