arXiv:2608.12683cs.ROcs.CV2026-08

让机器人主动找物体功能,比传统方法更聪明省力。

FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition

论文配图:FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition
图 1 · 摘自论文原文
  • 根据不确定性动态决定看哪里,边观察边规划。
  • 在多个数据集上表现最优,计算量减少33%。
  • 适合做智能机器人、自动驾驶的视觉决策系统。

具身智能体常需基于物体功能而非身份进行识别与交互,这就要求它们主动获取能揭示功能特征的观测信息。现有方法依赖固定视角,无法在功能线索被遮挡或不完整时决定该往何处观察。本文提出主动功能可及性定位新任务:智能体通过序列化探索场景,识别并精确定位满足功能查询的物体。为此,我们提出FUSE框架,结合显式不确定性驱动探索与学习到的近似规划器,高效选择有信息量的视角。此外,我们构建了基于Habitat的基准测试集以评估主动功能定位能力。实验表明,FUSE在无需先验知识的情况下达到最高观测性能,同时相较完全显式探索降低1.33倍计算开销,并在多种功能知识源下保持有效性。

原文摘要 · Abstract (English)

Embodied agents must often identify and interact with objects based on their function rather than their identity, requiring them to actively acquire observations that reveal discriminative functional evidence. Existing affordance grounding methods operate from fixed viewpoints and lack mechanisms for deciding where to look when functional cues are occluded or incomplete. We introduce Active Functional Affordance Grounding, a new task in which an agent sequentially explores a scene to identify and spatially ground an object satisfying a functional query. To address this problem, we propose FUSE, an adaptive semantic-geometric evidence acquisition framework that combines explicit uncertainty-driven exploration with a learned amortized planner to efficiently select informative viewpoints. We further introduce a Habitat-based benchmark for evaluating active functional grounding. Experiments show that FUSE achieves the highest observed non-oracle grounding performance while reducing computation by 1.33x relative to fully explicit exploration, and remains effective across multiple affordance knowledge sources.

具身智能主动感知视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。