arXiv:2605.23187cs.CVcs.RO2026-05

提出新基准IntentionNav,让智能体从隐含指令中推理目标物体。

IntentionNav: A Benchmark for Intent-Driven Object Navigation from Implicit Human Instruction

论文配图:IntentionNav: A Benchmark for Intent-Driven Object Navigation from Implicit Human Instruction
图 1 · 摘自论文原文
  • 基于隐式指令设计导航任务,不直接给出目标名称
  • 模型仅在24.9%场景成功到达目标,1米精度仅5.5%
  • 适合研究人机交互、意图理解与具身智能的学者

现有物体导航基准多指定明确类别(如微波炉或椅子),但真实场景中人类常给出间接指令,如“我需要加热这食物”或“房间太闷”。智能体需推断满足需求的物体,定位具体实例,并判断是否达成目标。本文提出意图为驱动的物体导航设置,并构建诊断性基准IntentionNav,包含500条自由文本意图、176个Isaac Sim场景和64类目标。每条意图以四种控制风格重写并标注四类意图模式,分离语言表达与语义线索类型,支持对目标推理、语言鲁棒性、邻域可达性和终点成功的独立分析。使用固定主动导航代理评估三个视觉语言模型,结果表明:模型在48.3%的回合中识别出正确目标,在68.7%中进入2米内邻域,但仅24.9%成功终止,1米精准落地成功率仅为5.5%。事件脚本类意图表现最佳(28.7%),物理状态与功能意图较低(19.2%和18.5%),说明隐式意图仍是具身搜索中目标选择、视觉验证与终端定位的核心瓶颈。

原文摘要 · Abstract (English)

Existing object navigation benchmarks usually tell an embodied agent which object category to find, such as microwave or chair. Human-facing embodied AI is often asked something less direct: "I need something to warm this food" or "the room feels stuffy." The agent must infer the object that can satisfy the need, find a scene-grounded instance, and decide whether the goal has been reached. We study this setting as intent-driven object navigation and introduce IntentionNav, a diagnostic benchmark for active object search from implicit human instructions. Each episode provides a free-text intent, RGB-D observations, and pose, but withholds the target object name. IntentionNav contains 500 intents over 176 Isaac Sim scenes and 64 target categories. Each intent is rewritten in four controlled instruction styles and annotated with one of four intent modes, separating surface phrasing from semantic cue type under matched geometry. This paired design supports analysis of target inference, language robustness, neighborhood reachability, and terminal success rather than only aggregate success. We evaluated three VLMs using a fixed active-navigation agent. Models identify the intended target in 48.3 percent of episodes and enter its 2 m neighborhood in 68.7 percent, but terminate successfully in only 24.9 percent and achieve grounded 1 m success in 5.5 percent. Success is highest for event-script intents (28.7 percent) and lower for physical-state and affordance intents (19.2 percent and 18.5 percent), showing that indirect human intent remains a bottleneck for target selection, visual verification, and terminal localization in active embodied search.

具身智能意图理解导航基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。