仅用一次视频演示,就能让机器人学会在不同形状物体上做相同动作。
Affordance Transfer Across Object Instances via Semantically Anchored Functional Map
- 通过语义锚点建立物体间功能区域的对应关系。
- 实现跨几何差异物体的交互区域精准迁移,计算开销小。
- 适合实际机器人感知-动作系统,结果可解释性强。
传统学习从示范(LfD)需要大量物理演示,耗时且难扩展。近期研究发现,机器人可通过人类视频提取交互线索,无需直接参与。然而核心挑战在于:如何将已演示的交互泛化到几何差异大但功能相似的不同物体实例。本文提出语义锚定功能映射(SemFM)框架,仅需一次视觉示范,即可在不同物体间转移功能属性。从图像重建粗略网格出发,方法识别物体间的语义对应功能区域,选取互斥的语义锚点,并通过功能映射在表面传播这些约束,获得密集且语义一致的对应关系。这使得示范中的交互区域能以轻量、可解释方式迁移到几何多样的物体上。在合成物体类别与真实机器人操作任务上的实验表明,该方法在计算成本较低的前提下实现高精度的属性迁移,适用于实际机器人感知-动作流水线。
原文摘要 · Abstract (English)
Traditional learning from demonstration (LfD) generally demands a cumbersome collection of physical demonstrations, which can be time-consuming and challenging to scale. Recent advances show that robots can instead learn from human videos by extracting interaction cues without direct robot involvement. However, a fundamental challenge remains: how to generalize demonstrated interactions across different object instances that share similar functionality but vary significantly in geometry. In this work, we propose \emph{Semantic Anchored Functional Maps} (SemFM), a framework for transferring affordances across objects from a single visual demonstration. Starting from a coarse mesh reconstructed from an image, our method identifies semantically corresponding functional regions between objects, selects mutually exclusive semantic anchors, and propagates these constraints over the surface using a functional map to obtain a dense, semantically consistent correspondence. This enables demonstrated interaction regions to be transferred across geometrically diverse objects in a lightweight and interpretable manner. Experiments on synthetic object categories and real-world robotic manipulation tasks show that our approach enables accurate affordance transfer with modest computational cost, making it well-suited for practical robotic perception-to-action pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。