arXiv:2511.08922cs.LGcs.AI2025-11

用扩散模型生成高质量动作,精准提升离线强化学习性能。

Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning

  • 根据动作优势值加权引导扩散模型,区分高价值与低价值动作。
  • 在D4RL上平均回报显著提升,AntMaze任务表现超越现有方法。
  • 适合追求高效探索与保守性平衡的离线强化学习研究者。

在离线强化学习中,分布外(OOD)动作导致的价值过估计严重限制策略性能。近期扩散模型因其强大的分布匹配能力被用于通过行为策略约束实现保守性。然而,现有方法对低质量数据集中冗余动作施加无差别正则化,造成过度保守,破坏扩散建模的表达力与效率平衡。为此,本文提出扩散策略的价值条件优化方法(DIVO),利用扩散模型生成高质量、广泛覆盖的分布内状态-动作样本,同时促进高效策略改进。具体地,DIVO引入二值加权机制,基于离线数据集中动作的优势值指导扩散模型训练,实现更精确的分布对齐,并选择性扩展高优势动作边界。在策略改进阶段,动态过滤扩散模型中高回报潜力的动作,有效引导学习策略向更优性能演进。该方法在离线RL中实现了保守性与探索性的关键平衡。我们在D4RL基准上评估DIVO,并与最先进基线对比。实验结果表明,DIVO在运动类任务中平均回报显著提升,在奖励稀疏的AntMaze领域表现优于现有方法。

原文摘要 · Abstract (English)

In offline reinforcement learning, value overestimation caused by out-of-distribution (OOD) actions significantly limits policy performance. Recently, diffusion models have been leveraged for their strong distribution-matching capabilities, enforcing conservatism through behavior policy constraints. However, existing methods often apply indiscriminate regularization to redundant actions in low-quality datasets, resulting in excessive conservatism and an imbalance between the expressiveness and efficiency of diffusion modeling. To address these issues, we propose DIffusion policies with Value-conditional Optimization (DIVO), a novel approach that leverages diffusion models to generate high-quality, broadly covered in-distribution state-action samples while facilitating efficient policy improvement. Specifically, DIVO introduces a binary-weighted mechanism that utilizes the advantage values of actions in the offline dataset to guide diffusion model training. This enables a more precise alignment with the dataset's distribution while selectively expanding the boundaries of high-advantage actions. During policy improvement, DIVO dynamically filters high-return-potential actions from the diffusion model, effectively guiding the learned policy toward better performance. This approach achieves a critical balance between conservatism and explorability in offline RL. We evaluate DIVO on the D4RL benchmark and compare it against state-of-the-art baselines. Empirical results demonstrate that DIVO achieves superior performance, delivering significant improvements in average returns across locomotion tasks and outperforming existing methods in the challenging AntMaze domain, where sparse rewards pose a major difficulty.

离线RL扩散模型策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。