arXiv:2409.05622cs.LG2024-09AAAI被引 12

用前向KL正则化直接对齐扩散策略与人类偏好,无需预设奖励函数。

Forward KL Regularized Preference Optimization for Aligning Diffusion Policies

  • 通过前向KL正则化在扩散策略中直接优化偏好对齐
  • 在MetaWorld和D4RL任务上优于现有最先进方法
  • 避免生成分布外动作,提升策略稳定性与安全性

扩散模型在序列决策任务中因强大的表达能力而表现卓越。学习扩散策略的核心挑战在于将策略输出对齐人类意图。以往方法依赖预定义奖励函数进行回报条件生成或基于强化学习的策略优化。本文提出一种新框架——前向KL正则化的偏好优化(Forward KL Regularized Preference Optimization),直接对齐扩散策略与偏好数据。首先从离线数据集训练扩散策略而不考虑偏好,随后通过直接偏好优化对策略进行对齐。在对齐阶段,我们将直接偏好学习引入扩散策略,并采用前向KL正则化防止生成分布外动作。我们在MetaWorld操控任务和D4RL基准上进行了广泛实验,结果表明该方法在偏好对齐方面表现更优,且超越现有最先进算法。

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable success in sequential decision-making by leveraging the highly expressive model capabilities in policy learning. A central problem for learning diffusion policies is to align the policy output with human intents in various tasks. To achieve this, previous methods conduct return-conditioned policy generation or Reinforcement Learning (RL)-based policy optimization, while they both rely on pre-defined reward functions. In this work, we propose a novel framework, Forward KL regularized Preference optimization for aligning Diffusion policies, to align the diffusion policy with preferences directly. We first train a diffusion policy from the offline dataset without considering the preference, and then align the policy to the preference data via direct preference optimization. During the alignment phase, we formulate direct preference learning in a diffusion policy, where the forward KL regularization is employed in preference optimization to avoid generating out-of-distribution actions. We conduct extensive experiments for MetaWorld manipulation and D4RL tasks. The results show our method exhibits superior alignment with preferences and outperforms previous state-of-the-art algorithms.

扩散模型策略对齐偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。