arXiv:2509.06233cs.ROcs.CV2025-09被引 7

一 shot学习3D物间交互功能,提升机器人泛化能力

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation

  • 用视觉大模型+点云融合语义与几何特征
  • 仅需1次样本即可准确识别物间功能关系
  • 适合需要快速适应新物体的机器人场景

物体功能接地是机器人操作的基础,它建立了感知与动作之间关于交互物体的关键联系。然而,以往工作主要关注单个物体的功能预测,忽视了现实中多数交互涉及物体对之间的关系。本文针对数据受限下的物对功能接地挑战,提出一种新颖的一次性3D物对功能学习方法。通过结合视觉基础模型的语义特征与点云表示的几何理解,我们的方法在新物体和新类别上具有良好的泛化能力。我们进一步将3D功能表示与大语言模型(LLMs)融合,显著提升了LLMs在生成任务特定约束函数时对物体交互的理解与推理能力。在3D物对功能接地和机器人操作上的实验表明,O$^3$Afford在准确性和泛化能力上均显著优于现有基线。

原文摘要 · Abstract (English)

Grounding object affordance is fundamental to robotic manipulation as it establishes the critical link between perception and action among interacting objects. However, prior works predominantly focus on predicting single-object affordance, overlooking the fact that most real-world interactions involve relationships between pairs of objects. In this work, we address the challenge of object-to-object affordance grounding under limited data contraints. Inspired by recent advances in few-shot learning with 2D vision foundation models, we propose a novel one-shot 3D object-to-object affordance learning approach for robotic manipulation. Semantic features from vision foundation models combined with point cloud representation for geometric understanding enable our one-shot learning pipeline to generalize effectively to novel objects and categories. We further integrate our 3D affordance representation with large language models (LLMs) for robotics manipulation, significantly enhancing LLMs' capability to comprehend and reason about object interactions when generating task-specific constraint functions. Our experiments on 3D object-to-object affordance grounding and robotic manipulation demonstrate that our O$^3$Afford significantly outperforms existing baselines in terms of both accuracy and generalization capability.

机器人操作3D感知少样本学习功能接地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。