arXiv:2511.22195cs.RO2025-11中稿 · IROS 2024被引 2

用3D关键点让机器人理解物体如何被使用,提升抓取精准度。

3D Affordance Keypoint Detection for Robotic Manipulation

  • 引入RGB与深度图像融合的3D关键点四元组,定位操作位置与方向。
  • 在分割与关键点检测任务中均显著优于现有模型。
  • 可直接指导机器人对未知物体完成可靠抓取操作。

本文提出一种基于3D关键点的新型具身功能感知方法,通过引入3D关键点四元组,结合RGB与深度图像信息,实现对物体操作位置、方向与作用范围的精准建模。该方法突破传统语义分割仅回答‘是什么’的局限,直接提供‘如何用’的执行指引。所提出的融合式具身关键点网络(FAKP-Net)在多个基准测试中,在具身分割与关键点检测任务上均取得显著领先。真实世界实验表明,该方法能有效指导机器人对未见过的物体完成稳定可靠的操控任务。

原文摘要 · Abstract (English)

This paper presents a novel approach for affordance-informed robotic manipulation by introducing 3D keypoints to enhance the understanding of object parts' functionality. The proposed approach provides direct information about what the potential use of objects is, as well as guidance on where and how a manipulator should engage, whereas conventional methods treat affordance detection as a semantic segmentation task, focusing solely on answering the what question. To address this gap, we propose a Fusion-based Affordance Keypoint Network (FAKP-Net) by introducing 3D keypoint quadruplet that harnesses the synergistic potential of RGB and Depth image to provide information on execution position, direction, and extent. Benchmark testing demonstrates that FAKP-Net outperforms existing models by significant margins in affordance segmentation task and keypoint detection task. Real-world experiments also showcase the reliability of our method in accomplishing manipulation tasks with previously unseen objects.

机器人抓取3D关键点具身感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。