arXiv:2601.06133cs.LGcs.AI2026-01综述被引 2

首次系统分析在线扩散策略强化学习算法,揭示其优劣与瓶颈。

A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control

  • 按策略优化机制分为四类:梯度、Q加权、邻近性与时间反向传播方法。
  • 在12个任务上验证,发现各方法在样本效率与可扩展性间存在根本权衡。
  • 适合关注机器人可控性与算法部署的科研人员参考。

扩散策略在机器人控制中展现出优越的多模态动作建模能力,但其与在线强化学习的结合仍面临根本性挑战,源于扩散模型训练目标与标准强化学习策略改进机制之间的不兼容。本文首次对面向可扩展机器人控制的在线扩散策略强化学习(Online DPRL)算法进行了全面综述与实证分析。我们提出一种新分类体系,将现有方法划分为四类:动作梯度、Q加权、基于邻近性与通过时间反向传播(BPTT)的方法。在统一的NVIDIA Isaac Lab基准下,涵盖12种不同机器人任务,系统评估了代表性算法在任务多样性、并行化能力、扩散步数可扩展性、跨本体泛化及环境鲁棒性五个维度的表现。分析揭示了各类算法固有的关键权衡,尤其在样本效率与可扩展性方面。同时,识别出制约在线DPRL实际部署的计算与算法瓶颈。基于此,我们为特定应用场景提供算法选择建议,并指明未来研究方向,推动更通用、可扩展的机器人学习系统发展。

原文摘要 · Abstract (English)

Diffusion policies have emerged as a powerful approach for robotic control, demonstrating superior expressiveness in modeling multimodal action distributions compared to conventional policy networks. However, their integration with online reinforcement learning remains challenging due to fundamental incompatibilities between diffusion model training objectives and standard RL policy improvement mechanisms. This paper presents the first comprehensive review and empirical analysis of current Online Diffusion Policy Reinforcement Learning (Online DPRL) algorithms for scalable robotic control systems. We propose a novel taxonomy that categorizes existing approaches into four distinct families--Action-Gradient, Q-Weighting, Proximity-Based, and Backpropagation Through Time (BPTT) methods--based on their policy improvement mechanisms. Through extensive experiments on a unified NVIDIA Isaac Lab benchmark encompassing 12 diverse robotic tasks, we systematically evaluate representative algorithms across five critical dimensions: task diversity, parallelization capability, diffusion step scalability, cross-embodiment generalization, and environmental robustness. Our analysis identifies key findings regarding the fundamental trade-offs inherent in each algorithmic family, particularly concerning sample efficiency and scalability. Furthermore, we reveal critical computational and algorithmic bottlenecks that currently limit the practical deployment of online DPRL. Based on these findings, we provide concrete guidelines for algorithm selection tailored to specific operational constraints and outline promising future research directions to advance the field toward more general and scalable robotic learning systems.

强化学习扩散模型机器人控制算法综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。