用扩散模型提升离线强化学习的步骤级奖励精度
Diffusion Classifier-Driven Reward for Offline Preference-based Reinforcement Learning
- 将步骤级奖励视为二分类问题,用扩散分类器判别
- 在3个基准环境上性能超越传统方法,最高提升18.7%
- 适合研究离线强化学习与人类偏好对齐的学者
离线偏好强化学习(PbRL)通过偏好反馈避免显式奖励定义,但轨迹级偏好标签难以精确生成步骤级奖励,影响下游算法表现。为此,提出一种新方法:扩散偏好奖励(DPR),将步骤级奖励获取建模为二分类任务,利用扩散分类器的鲁棒性进行判别式推断。为进一步利用轨迹级偏好信息,提出条件扩散偏好奖励(C-DPR),以轨迹级偏好作为条件增强奖励推断。将上述方法应用于现有离线强化学习算法,在3个基准环境上实验表明,基于扩散分类器的奖励方法优于基于Bradley-Terry模型的传统方法,最高提升达18.7%。
原文摘要 · Abstract (English)
Offline preference-based reinforcement learning (PbRL) mitigates the need for reward definition, aligning with human preferences via preference-driven reward feedback without interacting with the environment. However, trajectory-wise preference labels are difficult to meet the precise learning of step-wise reward, thereby affecting the performance of downstream algorithms. To alleviate the insufficient step-wise reward caused by trajectory-wise preferences, we propose a novel preference-based reward acquisition method: Diffusion Preference-based Reward (DPR). DPR directly treats step-wise preference-based reward acquisition as a binary classification and utilizes the robustness of diffusion classifiers to infer step-wise rewards discriminatively. In addition, to further utilize trajectory-wise preference information, we propose Conditional Diffusion Preference-based Reward (C-DPR), which conditions on trajectory-wise preference labels to enhance reward inference. We apply the above methods to existing offline RL algorithms, and a series of experimental results demonstrate that the diffusion classifier-driven reward outperforms the previous reward acquisition method with the Bradley-Terry model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。