arXiv:2603.19166cs.ROcs.AI2026-03

让机器人更准理解带距离的指令,比如‘往冰箱右边两米’。

Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation

  • 用多智能体分步解析语言指令,再融合结果
  • 在HM-EQA上比现有方法提升12.3%成功率
  • 适合需要精准空间定位的机器人任务

机器人与人类协作需将自然语言目标转化为可执行的物理动作。例如执行“往冰箱右侧两米处走”这类指令,需同时理解语义、空间关系和具体距离。尽管当前视觉语言模型(VLM)具备较强语义理解能力,但缺乏对物理空间中度量约束的显式推理。本文实证表明,主流基于VLM的指称消解方法在复杂度量-语义语言查询下表现不佳。为此,提出MAPG(多智能体概率指称),通过将语言指令分解为结构化子组件,分别调用VLM进行指称,并以概率方式融合输出,生成符合度量一致性的3D空间动作决策。在HM-EQA基准测试中,MAPG显著优于多个强基线。此外,构建新基准MAPG-Bench,专门评估度量-语义目标指称能力,弥补现有评估体系空白。还展示了真实机器人实验,在提供结构化场景表示时,MAPG可从仿真迁移到现实环境。

原文摘要 · Abstract (English)

Robots collaborating with humans must convert natural language goals into actionable, physically grounded decisions. For example, executing a command such as "go two meters to the right of the fridge" requires grounding semantic references, spatial relations, and metric constraints within a 3D scene. While recent vision language models (VLMs) demonstrate strong semantic grounding capabilities, they are not explicitly designed to reason about metric constraints in physically defined spaces. In this work, we empirically demonstrate that state-of-the-art VLM-based grounding approaches struggle with complex metric-semantic language queries. To address this limitation, we propose MAPG (Multi-Agent Probabilistic Grounding), an agentic framework that decomposes language queries into structured subcomponents and queries a VLM to ground each component. MAPG then probabilistically composes these grounded outputs to produce metrically consistent, actionable decisions in 3D space. We evaluate MAPG on the HM-EQA benchmark and show consistent performance improvements over strong baselines. Furthermore, we introduce a new benchmark, MAPG-Bench, specifically designed to evaluate metric-semantic goal grounding, addressing a gap in existing language grounding evaluations. We also present a real-world robot demonstration showing that MAPG transfers beyond simulation when a structured scene representation is available.

视觉语言导航多智能体空间推理机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。