arXiv:2604.02060cs.CVcs.RO2026-04被引 1

让机器人在多个相似物体中根据任务意图选对工具

CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects

论文配图:CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
图 1 · 摘自论文原文
  • 用语言意图引导,在多物体点云中定位正确功能对象
  • 构建首个含8.8万组问答的混淆对基准数据集
  • 适合研究具身智能与人机交互的学者参考

当被要求“切蛋糕”时,机器人需从刀和剪刀中选择刀具,尽管二者均具备切割功能。现实场景中,多个物体可能共享相同功能,但仅有一个在特定任务下适用,这类情况称为混淆对。现有3D功能识别方法多忽略此挑战,通常评估孤立单个物体且查询中显式提供类别名。本文提出意图驱动的混淆功能定位新任务,要求在多物体点云中基于隐式自然语言意图预测正确对象的逐点功能掩码。为此构建了首个聚焦隐式意图的基准数据集CompassAD,包含30组混淆物体对、16种功能类型、6,422种组合及88,000+组问答对。同时提出CompassNet框架,引入实例边界交叉注入(ICI)模块,限制语言-几何对齐范围以防止跨对象语义泄露;采用双层对比精炼(BCR)模块,在几何群组与点级实现双重区分,强化目标与混淆表面间的差异。大量实验表明该方法在已见与未见查询上均达领先性能,且在机械臂上的部署验证了其在真实复杂场景中抓取的有效迁移能力。

原文摘要 · Abstract (English)

When told to "cut the cake," a robot must choose the knife over nearby scissors, despite both objects affording the same cutting function. In real-world scenes, multiple objects may share identical affordances, yet only one is appropriate under the given task context. We call such cases confusing pairs. However, existing 3D affordance methods largely sidestep this challenge by evaluating isolated single objects, often with explicit category names provided in the query. We formalize Intent-Driven Confusable Affordance Grounding, a new 3D affordance setting that requires predicting a per-point affordance mask on the correct object within a multi-object point cloud, conditioned on implicit natural language intent. To study this problem, we construct CompassAD, the first benchmark centered on implicit intent in confusing multi-object compositions. It comprises 30 confusing object pairs spanning 16 affordance types, 6,422 compositions, and 88K+ query-answer pairs. Furthermore, we propose CompassNet, a framework that incorporates two dedicated modules tailored to this task. Instance-bounded Cross Injection (ICI) constrains language-geometry alignment within object boundaries to prevent cross-object semantic leakage. Bi-level Contrastive Refinement (BCR) enforces discrimination at both geometric-group and point levels, sharpening distinctions between target and confusable surfaces. Extensive experiments demonstrate state-of-the-art results on both seen and unseen queries, and deployment on a robotic manipulator confirms effective transfer to real-world grasping in confusing multi-object compositions.

3D感知机器人意图理解功能定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。