arXiv:2506.07822cs.LGcs.AI2025-06被引 3

让离线强化学习的扩散规划更快更优,一步生成高奖励动作轨迹。

Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation

  • 将奖励优化融入一致性蒸馏,解耦训练且无需噪声信号。
  • 单步采样下比当前最佳方法提升9.7%奖励,推理速度加快142倍。
  • 适合需要高效高质决策的长周期规划任务,如机器人控制。

尽管扩散模型在决策任务中表现优异,但其推理速度慢仍是关键瓶颈。虽然一致性模型提供了潜在解决方案,但现有决策应用要么在行为克隆下产生次优示范,要么依赖复杂的演员-评论家框架中多网络并行训练。本文提出一种新型一致性蒸馏方法,直接将奖励优化整合进蒸馏过程。该方法实现单步采样,通过解耦训练和无噪声奖励信号生成更高奖励的动作轨迹。在Gym MuJoCo、FrankaKitchen及长周期规划基准上的实证评估表明,本方法相比先前最优结果提升9.7%,推理速度较扩散模型最高提速142倍。

原文摘要 · Abstract (English)

Although diffusion models have achieved strong results in decision-making tasks, their slow inference speed remains a key limitation. While consistency models offer a potential solution, existing applications to decision-making either struggle with suboptimal demonstrations under behavior cloning or rely on complex concurrent training of multiple networks under the actor-critic framework. In this work, we propose a novel approach to consistency distillation for offline reinforcement learning that directly incorporates reward optimization into the distillation process. Our method achieves single-step sampling while generating higher-reward action trajectories through decoupled training and noise-free reward signals. Empirical evaluations on the Gym MuJoCo, FrankaKitchen, and long horizon planning benchmarks demonstrate that our approach can achieve a 9.7% improvement over previous state-of-the-art while offering up to 142x speedup over diffusion counterparts in inference time.

扩散模型强化学习加速推理轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。