提出新算法让离线偏好强化学习更稳定高效。
Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning
- 将偏好学习建模为策略与模型的博弈,实现可计算的保守性约束。
- 在连续控制任务中表现接近当前最优方法,且无需复杂置信集构造。
- 适合追求理论保障与实际性能平衡的研究者和工程师。
本文研究离线偏好强化学习(PbRL),即基于预收集的轨迹对偏好反馈进行学习。尽管离线PbRL已展现出显著的实证成功,但现有理论方法在不确定性下难以保证保守性,且依赖计算上不可行的置信集构造。为此,我们提出对抗式偏好策略优化(APPO),一种计算高效的离线PbRL算法,可在不依赖显式置信集的情况下保证样本复杂度界。通过将PbRL建模为策略与模型之间的双人博弈,该方法以可计算方式实现保守性。在标准函数近似与有界轨迹集中性假设下,我们推导出样本复杂度界。据我们所知,APPO是首个同时具备统计效率与实际可用性的离线PbRL算法。在连续控制任务上的实验表明,APPO能有效从复杂数据集中学习,性能与现有最先进方法相当。
原文摘要 · Abstract (English)
In this paper, we study offline preference-based reinforcement learning (PbRL), where learning is based on pre-collected preference feedback over pairs of trajectories. While offline PbRL has demonstrated remarkable empirical success, existing theoretical approaches face challenges in ensuring conservatism under uncertainty, requiring computationally intractable confidence set constructions. We address this limitation by proposing Adversarial Preference-based Policy Optimization (APPO), a computationally efficient algorithm for offline PbRL that guarantees sample complexity bounds without relying on explicit confidence sets. By framing PbRL as a two-player game between a policy and a model, our approach enforces conservatism in a tractable manner. Using standard assumptions on function approximation and bounded trajectory concentrability, we derive a sample complexity bound. To our knowledge, APPO is the first offline PbRL algorithm to offer both statistical efficiency and practical applicability. Experimental results on continuous control tasks demonstrate that APPO effectively learns from complex datasets, showing comparable performance with existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。