arXiv:2506.18779cs.RO2025-06被引 2

用扩散模型生成可变形物体的多种目标形状,提升机器人操作适应性。

DiffDef: A Diffusion Model for Generating Multimodal Goal Shapes From Demonstrations for Deformable Object Manipulation

  • 基于扩散模型学习可行目标形状的分布,而非单一确定结果。
  • 在仿真与真实机器人上验证,显著提升多目标场景任务成功率。
  • 适用于制造、手术等需灵活应对多种目标形态的复杂操作场景。

可变形物体操作是众多机器人应用中的关键技术。现有方法通常依赖不切实际的目标形状获取方式,如领域知识工程或手动调整,且普遍假设单一确定性目标,难以处理现实任务中常见的多模态目标情况——即多个不同形状均能成功完成任务。本文提出 DiffDef,一种基于扩散模型的神经网络,能够学习可行目标形状的分布,而非预测单一结果。该方法可生成多样化的目标配置,避免确定性预测中的模式平均化缺陷。我们在受制造与手术任务启发的多个可变形物体操作任务上进行了评估,涵盖仿真环境及两个物理机器人平台:da Vinci Research Kit(dVRK)和双臂KUKA系统。实验结果表明,DiffDef能有效捕捉多模态目标分布,并在实际机器人场景中显著提升任务表现。

原文摘要 · Abstract (English)

Deformable object manipulation is a key capability in many robotic applications. A promising paradigm for this problem is shape servoing, which aims to control deformable objects toward desired goal shapes. However, existing approaches typically rely on impractical goal-shape acquisition methods, such as domain-knowledge engineering or manual manipulation. Moreover, prior methods generally assume a single deterministic goal and fail to handle multimodal goal settings, a common scenario in many real-world tasks where multiple distinct goal shapes can all lead to successful task completion. In this paper, we introduce DiffDef, a novel neural network that uses a diffusion model to learn a distribution of feasible goal shapes rather than predicting a single deterministic outcome. This allows DiffDef to generate diverse goal configurations while avoiding the mode-averaging artifacts common in deterministic predictors. We evaluate our method on several deformable manipulation tasks inspired by manufacturing and surgical applications, both in simulation and on two physical robotic platforms: the da Vinci Research Kit (dVRK) and a bimanual KUKA-based robotic system. The results demonstrate that DiffDef effectively captures multimodal goal distributions and significantly improves task performance in practical robotic settings. Website: sites.google.com/view/diffdef.

可变形操作扩散模型多模态生成机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。