arXiv:2608.13605cs.AIcs.RO2026-08

机器人通过主动移动观察来消除目标歧义,提升理解与决策能力。

Active Perception for Embodied Disambiguation

论文配图:Active Perception for Embodied Disambiguation
图 1 · 摘自论文原文
  • 用主动观察获取视觉信息,替代仅依赖用户提问的被动方式。
  • 在真实机器人上验证,显著减少用户澄清需求并提高任务成功率。
  • 适合需要自主感知与交互的智能机器人应用,如家庭服务或导航。

自然语言为机器人提供灵活的任务接口,但具身环境中的目标歧义不仅源于用户意图,也可能因当前观测中缺少相关物理证据所致。现有交互消歧方法主要依赖向用户提问获取更多信息,而遮挡、视角受限、文字无法识别及目标未被观测等情况,要求机器人主动改变自身观测位置。本文提出一种面向具身目标消歧的主动感知框架,以主动观察作为信息获取的核心,并利用视觉-语言模型基于累积的视觉证据与交互信息,判断是否继续观察、请求澄清或完成目标选择。主动观察不仅能直接恢复缺失的判别性信息,还可揭示物体名称、标签与语义属性,从而在必要时提升用户澄清效率。真实机器人实验表明,该框架将物理信息获取与用户意图澄清整合进统一的具身消歧流程中。

原文摘要 · Abstract (English)

Natural language provides robots with a flexible task interface, but target ambiguity in embodied environments arises not only from user intent; it can also result from missing taskrelevant physical evidence in the current observation. Existing interactive disambiguation methods primarily obtain additional information by asking the user, whereas occlusion, restricted viewpoints, unreadable text, and unobserved targets require the robot to actively change its observation. We propose an active-perception framework for embodied target disambiguation that uses active observation as the backbone for information acquisition and uses a vision-language model to decide, on the basis of accumulated visual evidence and interaction information, whether to continue observing, request clarification, or complete target selection. Active observation can both directly recover missing discriminative evidence and reveal object names, labels, and semantic attributes, thereby improving user clarification when it remains necessary. Real-robot experiments show that the framework combines physical information acquisition and userintent clarification within a unified embodied disambiguation process.

具身智能主动感知人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。