扩散模型赋能强化学习,实现多模态规划与稳定训练。
Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development
- 构建双轴分类体系,梳理扩散模型在RL中的功能与技术路径。
- 覆盖单/多智能体场景,形成可复用的扩散-强化学习集成框架。
- 适合关注生成式强化学习、轨迹规划的研究者参考。
扩散模型(DMs)作为主流生成模型,为强化学习(RL)带来多模态表达、稳定训练和轨迹级规划等优势。本综述系统梳理了基于扩散模型的强化学习研究进展。首先概述强化学习的挑战,并介绍扩散模型的基础原理,分析其如何融入强化学习框架以应对核心问题。提出一个双轴分类体系:功能导向维度明确扩散模型在强化学习流程中的角色;技术导向维度区分在线与离线学习范式下的实现方式。进一步从单智能体拓展至多智能体领域,总结多种扩散-强化学习融合架构并强调其实际应用价值。还列举了扩散模型在多个领域的成功应用案例,讨论当前方法的开放问题,并指出未来关键研究方向。最后,归纳未来发展趋势。相关论文与资源持续维护于GitHub仓库(https://github.com/ChangfuXu/D4RL-FTD)。
原文摘要 · Abstract (English)
Diffusion Models (DMs), as a leading class of generative models, offer key advantages for reinforcement learning (RL), including multi-modal expressiveness, stable training, and trajectory-level planning. This survey delivers a comprehensive and up-to-date synthesis of diffusion-based RL. We first provide an overview of RL, highlighting its challenges, and then introduce the fundamental concepts of DMs, investigating how they are integrated into RL frameworks to address key challenges in this research field. We establish a dual-axis taxonomy that organizes the field along two orthogonal dimensions: a function-oriented taxonomy that clarifies the roles DMs play within the RL pipeline, and a technique-oriented taxonomy that situates implementations across online versus offline learning regimes. We also provide a comprehensive examination of this progression from single-agent to multi-agent domains, thereby forming several frameworks for DM-RL integration and highlighting their practical utility. Furthermore, we outline several categories of successful applications of diffusion-based RL across diverse domains, discuss open research issues of current methodologies, and highlight key directions for future research to advance the field. Finally, we summarize the survey to identify promising future development directions. We are actively maintaining a GitHub repository (https://github.com/ChangfuXu/D4RL-FTD) for papers and other related resources to apply DMs for RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。