用信任域机制让扩散模型在并行强化学习中稳定训练
Trust-Region Diffusion Policies for Massively Parallel On-Policy RL

- 引入信任域约束,控制扩散路径的KL散度以稳定训练
- 在73个任务上表现优于或媲美现有基线,尤其在人形控制中优势明显
- 适合需要高稳定性与复杂策略的并行强化学习场景
大规模并行模拟已成为开发可部署强化学习策略的标准框架,但多数方法仍依赖简单的高斯策略参数化。扩散模型具备更强表达能力,在复杂控制问题上表现优异,然而现有基于扩散模型的强化学习方法多面向离线或非策略训练。本文探讨扩散策略能否在大规模并行、在线策略设置下有效训练。为此,提出信任域扩散策略(TruDi),实现基于扩散模型的在线强化学习。该设置极具挑战性,因数据分布随更新快速变化,难以维持复杂策略的稳定训练。TruDi通过引入信任域优化规则,对整个扩散轨迹施加KL散度约束以提升稳定性。我们在涵盖73个任务的4个大规模并行强化学习基准上评估了TruDi。结果表明,其在标准任务上持续优于或媲美强基线,在更具挑战性的人形控制任务上取得显著提升,为大规模并行在线强化学习建立了新基准。
原文摘要 · Abstract (English)
Reinforcement learning with massively parallel simulations has become a standard framework for developing robust, deployable policies; however, most existing approaches still rely on simple Gaussian policy parameterizations. Diffusion models provide a more expressive policy class and have shown strong performance on challenging control problems, yet most diffusion-based RL methods are designed for offline or off-policy training. In this work, we ask whether diffusion policies can be trained effectively in the massively parallel, on-policy regime. To this end, we introduce Trust-region Diffusion Policies (TruDi), which enables diffusion policies for on-policy RL with massively parallel simulations. This setting is particularly challenging because the data distribution changes quickly across updates, making stable training with complex policies difficult. TruDi addresses this by integrating a trust-region optimization rule to enforce a KL-divergence constraint over the entire diffusion trajectory. Empirically, we evaluate TruDi on a diverse set of 4 massively parallel RL benchmarks comprising a total of 73 tasks. Across these tasks, TruDi consistently outperforms or is on-par with strong baselines on standard tasks and achieves clear gains on more challenging humanoid control tasks, establishing a strong new baseline for massively parallel on-policy RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。