arXiv:2504.08438cs.ROstat.ML2025-04综述被引 75

扩散模型让机器人抓取与规划更智能,突破传统方法局限。

Diffusion Models for Robotic Manipulation: A Survey

  • 用概率框架建模多模态动作分布,适应复杂环境
  • 在抓取与轨迹规划任务中提升成功率与泛化能力
  • 适合做视觉引导的机器人学习,尤其数据少时

扩散生成模型在图像与视频生成等视觉领域表现卓越,近期也展现出在机器人操作中的巨大潜力。这类模型基于概率框架,擅长建模多模态分布,并对高维输入输出空间具有强鲁棒性。本文全面综述了当前最先进的扩散模型在机器人操作中的应用,涵盖抓取学习、轨迹规划与数据增强。其中,场景与图像增强方法融合了机器人与计算机视觉,提升了视觉任务的泛化能力与数据稀缺场景下的性能。论文还介绍了扩散模型的两大主流架构及其与模仿学习、强化学习的结合方式,分析了常见网络结构与基准测试集,并指出当前方法的优势与挑战。

原文摘要 · Abstract (English)

Diffusion generative models have demonstrated remarkable success in visual domains such as image and video generation. They have also recently emerged as a promising approach in robotics, especially in robot manipulations. Diffusion models leverage a probabilistic framework, and they stand out with their ability to model multi-modal distributions and their robustness to high-dimensional input and output spaces. This survey provides a comprehensive review of state-of-the-art diffusion models in robotic manipulation, including grasp learning, trajectory planning, and data augmentation. Diffusion models for scene and image augmentation lie at the intersection of robotics and computer vision for vision-based tasks to enhance generalizability and data scarcity. This paper also presents the two main frameworks of diffusion models and their integration with imitation learning and reinforcement learning. In addition, it discusses the common architectures and benchmarks and points out the challenges and advantages of current state-of-the-art diffusion-based methods.

扩散模型机器人操作数据增强多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。