arXiv:2504.00280cs.AI2025-04

扩散策略在动态任务中表现更优,适应性强。

Exploration and Adaptation in Non-Stationary Tasks with Diffusion Policies

  • 用迭代去噪生成连续动作序列,从视觉输入中学习
  • 在Procgen和PointMaze上均优于PPO、DQN,奖励更高且波动小
  • 适合机器人装配、自动驾驶等动态场景的实时适应

本文研究扩散策略在非平稳、基于视觉的强化学习环境中的应用,聚焦于任务动态与目标随时间变化的场景。研究针对真实世界中如机器人装配线和自主导航等实际挑战,要求智能体从高维视觉输入中自适应调整控制策略。采用扩散策略——通过迭代随机去噪优化潜在动作表示——在Procgen和PointMaze等基准环境中进行测试。实验表明,尽管计算开销增加,扩散策略仍持续优于标准RL方法(如PPO和DQN),在平均奖励和最大奖励上均有提升,且方差更低。结果证明该方法能在持续变化的条件下生成连贯且情境相关的动作序列,同时揭示了在极端非平稳性下进一步优化的空间。

原文摘要 · Abstract (English)

This paper investigates the application of Diffusion Policy in non-stationary, vision-based RL settings, specifically targeting environments where task dynamics and objectives evolve over time. Our work is grounded in practical challenges encountered in dynamic real-world scenarios such as robotics assembly lines and autonomous navigation, where agents must adapt control strategies from high-dimensional visual inputs. We apply Diffusion Policy -- which leverages iterative stochastic denoising to refine latent action representations-to benchmark environments including Procgen and PointMaze. Our experiments demonstrate that, despite increased computational demands, Diffusion Policy consistently outperforms standard RL methods such as PPO and DQN, achieving higher mean and maximum rewards with reduced variability. These findings underscore the approach's capability to generate coherent, contextually relevant action sequences in continuously shifting conditions, while also highlighting areas for further improvement in handling extreme non-stationarity.

扩散策略强化学习动态任务视觉控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。