让智能体通过对话实现长时程目标导航,解决真实指令模糊问题。
VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs
- 设计可主动提问的导航系统,结合语言与移动决策。
- 构建41000+条带对话的长时程导航轨迹数据集。
- 适合研究交互式机器人导航与人机协作的学者。
现有具身导航任务多基于清晰明确的指令,如指令跟随或物品搜索。现实中指令常不完整,需通过互动澄清歧义、理解用户意图。为此,我们提出交互式实例目标导航(IIGN),在实例目标导航(IGN)基础上,允许智能体以自然语言自由向裁判提问,同时在环境中寻找特定目标。IIGN要求智能体输出语言与导航双重响应,实现在移动中持续交互。为此,我们构建了VL-LN Bench基准,包含自动化数据采集管道和超过41,000条长时程对话增强轨迹,配套自动评估协议及专用问答裁判。实验发现两大瓶颈:长时程探索能力不足,以及在同类干扰项中精准定位目标实例的能力弱。尽管主动对话部分缓解了这些问题,当前模型仍远低于人类表现。消融实验验证了我们数据生成管道的价值,并证明所提裁判提供可扩展的辅助效果接近人工支持,表明VL-LN Bench是对话驱动具身导航的实用测试平台。
原文摘要 · Abstract (English)
In most existing embodied navigation tasks, instructions are well-defined and unambiguous, such as instruction following and object searching. Under this idealized setting, agents are required solely to produce effective navigation outputs conditioned on vision and language (VL) inputs. Real-world instructions, however, are often underspecified and require interaction to resolve ambiguity and infer user intent. To bridge this gap, we propose Interactive Instance Goal Navigation (IIGN), which extends Instance Goal Navigation (IGN) by allowing agents to freely consult an oracle in natural language while searching for a specific instance. IIGN requires agents to produce both Language and Navigation (LN) outputs, enabling interaction while moving in the environment. To support this task, we introduce VL-LN Bench, a benchmark with an automated data collection pipeline and over 41k collected long-horizon dialog-augmented trajectories for training, alongside an automatic evaluation protocol paired with a dedicated oracle for answering agent queries. Experiments reveal two core bottlenecks of IIGN: long-horizon exploration and fine-grained grounding of textual information to the correct instance among same-category distractors. Although active dialog partially alleviates these challenges, current models still lag far behind human performance. Further ablations validate the value of the data generated by our pipeline and show that the proposed oracle provides scalable assistance comparable to human support, proving VL-LN Bench as a practical testbed for dialog-enabled embodied navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。