arXiv:2502.12198cs.LGcs.AI2025-02被引 1

提出统一框架,提升扩散模型在控制任务中的奖励最大化能力。

Maximize Your Diffusion: A Study into Reward Maximization and Alignment for Diffusion-based Control

  • 融合四种微调方法,构建统一的奖励优化框架。
  • 在离线强化学习中实现多任务控制性能提升。
  • 适合研究扩散模型决策与对齐的学者参考。

基于扩散的规划、学习与控制方法为强大且表达力强的决策方案提供了新路径。尽管近年来已有诸多改进,但现有方法在决策过程中通用的奖励最大化策略方面仍显不足。本文研究了控制应用中微调方法的扩展,具体探讨了四种微调方法的延伸与设计选择:通过强化学习进行奖励对齐、直接偏好优化、监督微调以及级联扩散。我们优化其使用方式,将这些独立努力整合为一个统一范式。实验证明该方法在离线强化学习设置下具有实用性,并在丰富多样的控制任务中展现出显著性能提升。

原文摘要 · Abstract (English)

Diffusion-based planning, learning, and control methods present a promising branch of powerful and expressive decision-making solutions. Given the growing interest, such methods have undergone numerous refinements over the past years. However, despite these advancements, existing methods are limited in their investigations regarding general methods for reward maximization within the decision-making process. In this work, we study extensions of fine-tuning approaches for control applications. Specifically, we explore extensions and various design choices for four fine-tuning approaches: reward alignment through reinforcement learning, direct preference optimization, supervised fine-tuning, and cascading diffusion. We optimize their usage to merge these independent efforts into one unified paradigm. We show the utility of such propositions in offline RL settings and demonstrate empirical improvements over a rich array of control tasks.

扩散模型强化学习控制奖励对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。