arXiv:2412.05268cs.ROcs.CV2024-12被引 42

通过单次演示实现跨类别3D语义匹配,提升机器人操作泛化能力

DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single Demo

  • 基于多视角2D特征投影与3D网络优化,生成物体顶点特征
  • 在跨类别物体间实现43.5%的匹配性能提升
  • 适合需要少样本泛化的机器人抓取与数字资产外观迁移场景

3D稠密对应可增强机器人操作能力,实现从一个物体到未见对象的空间、功能和动态信息泛化。相比形状对应,语义对应在不同类别间泛化更有效。为此,我们提出DenseMatcher,一种可在结构相似的野外物体间计算3D对应的方法。该方法首先将多视角2D特征投影到网格上,并通过3D网络精炼顶点特征,随后利用函数映射(functional map)寻找稠密对应。此外,我们构建了首个包含多样类别彩色物体网格的3D匹配数据集。实验表明,DenseMatcher相比先前3D匹配基线显著提升43.5%。我们展示了其在下游任务中的有效性:(i) 在机器人操作中,仅需观察一次示范即可实现长时序复杂任务的跨实例、跨类别泛化;(ii) 在零样本颜色映射中,可将外观在几何相关物体间迁移。

原文摘要 · Abstract (English)

Dense 3D correspondence can enhance robotic manipulation by enabling the generalization of spatial, functional, and dynamic information from one object to an unseen counterpart. Compared to shape correspondence, semantic correspondence is more effective in generalizing across different object categories. To this end, we present DenseMatcher, a method capable of computing 3D correspondences between in-the-wild objects that share similar structures. DenseMatcher first computes vertex features by projecting multiview 2D features onto meshes and refining them with a 3D network, and subsequently finds dense correspondences with the obtained features using functional map. In addition, we craft the first 3D matching dataset that contains colored object meshes across diverse categories. In our experiments, we show that DenseMatcher significantly outperforms prior 3D matching baselines by 43.5%. We demonstrate the downstream effectiveness of DenseMatcher in (i) robotic manipulation, where it achieves cross-instance and cross-category generalization on long-horizon complex manipulation tasks from observing only one demo; (ii) zero-shot color mapping between digital assets, where appearance can be transferred between different objects with relatable geometry.

3D匹配机器人操作少样本学习语义对应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。