arXiv:2410.23039cs.ROcs.CV2024-10CoRL被引 10

用注意力机制建模3D场景点间关系,实现一次演示的灵巧抓取。

Neural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping

  • 通过跨点注意力机制聚合全局语义特征,替代传统点特征表示
  • 仅需少量点云数据训练,在真实机器人上抓取成功率显著提升
  • 适合缺乏标注数据的灵巧操作任务,尤其适用于新场景快速适配

一次性将灵巧抓取能力迁移到存在物体与环境变化的新场景仍具挑战。尽管大视觉模型提取的特征场能建立3D场景间的语义对应,但其特征为点级且局限于物体表面,难以建模手-物交互的复杂语义分布。本文提出神经注意力场(Neural Attention Field),通过建模点间相关性,而非单个点特征,实现3D空间中语义感知的密集特征表示。核心是使用Transformer解码器计算任意3D查询点与所有场景点之间的交叉注意力,并以注意力加权方式聚合特征。我们进一步设计无手部示范的自监督训练框架,仅需少量3D点云即可训练。训练后,该注意力场可应用于新场景,实现基于一次演示的语义感知灵巧抓取。实验表明,该方法通过引导末端执行器聚焦任务相关区域,改善了优化景观,相比基于特征场的方法,在真实机器人上抓取成功率有显著提升。

原文摘要 · Abstract (English)

One-shot transfer of dexterous grasps to novel scenes with object and context variations has been a challenging problem. While distilled feature fields from large vision models have enabled semantic correspondences across 3D scenes, their features are point-based and restricted to object surfaces, limiting their capability of modeling complex semantic feature distributions for hand-object interactions. In this work, we propose the \textit{neural attention field} for representing semantic-aware dense feature fields in the 3D space by modeling inter-point relevance instead of individual point features. Core to it is a transformer decoder that computes the cross-attention between any 3D query point with all the scene points, and provides the query point feature with an attention-based aggregation. We further propose a self-supervised framework for training the transformer decoder from only a few 3D pointclouds without hand demonstrations. Post-training, the attention field can be applied to novel scenes for semantics-aware dexterous grasping from one-shot demonstration. Experiments show that our method provides better optimization landscapes by encouraging the end-effector to focus on task-relevant scene regions, resulting in significant improvements in success rates on real robots compared with the feature-field-based methods.

灵巧抓取注意力机制3D场景理解零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。