arXiv:2412.03142cs.RO2024-12CVPR被引 49

用可迁移的交互先验提升扩散策略在新物体上的泛化能力

AffordDP: Generalizable Diffusion Policy with Transferable Affordance

  • 通过3D接触点和后接触轨迹建模物体交互先验
  • 在仿真与真实环境上均实现对未见类别和实例的成功泛化
  • 适合需要跨类别鲁棒操作的机器人任务

基于扩散的策略在机器人操作任务中表现优异,但在分布外场景下表现不佳。现有方法虽改进了视觉特征编码,但泛化能力通常仅限于外观相似的同类物体。本文提出一种可迁移交互先验的扩散策略(AffordDP),旨在实现跨新类别的通用操作。AffordDP通过3D接触点和后接触轨迹建模静态与动态交互信息,利用基础视觉模型与点云配准技术,估计6维变换矩阵,将领域内学习到的交互先验迁移到未见物体上。更重要的是,在扩散采样过程中引入交互先验引导,使生成动作逐步趋向目标操作,同时保持在动作空间流形内。在仿真与真实世界环境中实验表明,AffordDP显著优于以往扩散方法,可在其他方法失败的情况下成功泛化至未见实例与类别。

原文摘要 · Abstract (English)

Diffusion-based policies have shown impressive performance in robotic manipulation tasks while struggling with out-of-domain distributions. Recent efforts attempted to enhance generalization by improving the visual feature encoding for diffusion policy. However, their generalization is typically limited to the same category with similar appearances. Our key insight is that leveraging affordances--manipulation priors that define "where" and "how" an agent interacts with an object--can substantially enhance generalization to entirely unseen object instances and categories. We introduce the Diffusion Policy with transferable Affordance (AffordDP), designed for generalizable manipulation across novel categories. AffordDP models affordances through 3D contact points and post-contact trajectories, capturing the essential static and dynamic information for complex tasks. The transferable affordance from in-domain data to unseen objects is achieved by estimating a 6D transformation matrix using foundational vision models and point cloud registration techniques. More importantly, we incorporate affordance guidance during diffusion sampling that can refine action sequence generation. This guidance directs the generated action to gradually move towards the desired manipulation for unseen objects while keeping the generated action within the manifold of action space. Experimental results from both simulated and real-world environments demonstrate that AffordDP consistently outperforms previous diffusion-based methods, successfully generalizing to unseen instances and categories where others fail.

扩散模型机器人操作泛化能力交互先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。