通过解耦点云扩散实现高精度物体摆放,提升泛化能力。
Disentangled Point Diffusion for Precise Object Placement
- 分层解耦设计:先全局预测位置分布,再局部精调几何与姿态
- 在仿真和真实工业插入任务中实现最先进精度,误差显著低于基线
- 支持柔性物体(如布料)摆放,突破刚体假设限制
机器人操作中的学习演示已展现出强大潜力,但端到端策略在面对新物体形状时泛化能力差且精度不足。本文提出TAX-DPD框架,采用目标中心化建模思路,通过新型前馈密集高斯混合模型生成全局放置空间的稠密先验,并引入解耦点云扩散模块,分别处理物体几何与放置姿态,实现精细局部推理。实验表明,该方法在模拟与真实世界高精度工业插入任务中均达到当前最优性能,且在刚性物体摆放上优于基于SE(3)的扩散方法。此外,在模拟布料悬挂任务中也展现良好表现,说明其可放宽对物体刚性的假设。
原文摘要 · Abstract (English)
Recent advances in robotic manipulation have highlighted the effectiveness of learning from demonstration. However, while end-to-end policies excel in expressivity and flexibility, they struggle both in generalizing to novel object geometries and in attaining a high degree of precision. An alternative, object-centric approach frames the task as predicting the placement pose of the target object, providing a modular decomposition of the problem. Building on this goal-prediction paradigm, we propose TAX-DPD, a hierarchical, disentangled point diffusion framework that achieves state-of-the-art performance in placement precision, multi-modal coverage, and generalization to variations in object geometries and scene configurations. We model global scene-level placements through a novel feed-forward Dense Gaussian Mixture Model (GMM) that yields a spatially dense prior over global placements; we then model the local object-level configuration through a novel disentangled point cloud diffusion module that separately diffuses the object geometry and the placement frame, enabling precise local geometric reasoning. Interestingly, we demonstrate that our point cloud diffusion achieves substantially higher accuracy than a prior approach based on SE(3)-diffusion, even in the context of rigid object placement. We validate our approach across a suite of challenging tasks in simulation and in the real-world on high-precision industrial insertion tasks. Furthermore, we present results on a cloth-hanging task in simulation, indicating that our framework can further relax assumptions on object rigidity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。