让导航员与远程引导者通过多轮对话协作找路,强调实时位置理解。
DialNav: Multi-turn Dialog Navigation with a Remote Guide
- 构建多轮对话导航框架,引导者需推断导航者位置
- 发布真实环境下的真人对话与路径数据集 RAIN
- 提供完整评测体系,适合研究具身对话的学者
我们提出 DialNav,一种新型具身对话协作导航任务:导航代理(Navigator)与远程引导者(Guide)通过多轮对话抵达目标位置。与以往工作不同,DialNav 强调全面评估,要求引导者推断导航者的实际位置,使沟通成为任务成功的关键。为此,我们收集并发布了远程导航协助数据集(RAIN),包含在逼真环境中的人类-人类对话与导航轨迹。我们设计了涵盖导航与对话的综合评测基准,并开展大量实验,分析不同导航与引导模型的影响。本文揭示关键挑战,并公开数据集、代码与评估框架,以推动具身对话领域的研究。
原文摘要 · Abstract (English)
We introduce DialNav, a novel collaborative embodied dialog task, where a navigation agent (Navigator) and a remote guide (Guide) engage in multi-turn dialog to reach a goal location. Unlike prior work, DialNav aims for holistic evaluation and requires the Guide to infer the Navigator's location, making communication essential for task success. To support this task, we collect and release the Remote Assistance in Navigation (RAIN) dataset, human-human dialog paired with navigation trajectories in photorealistic environments. We design a comprehensive benchmark to evaluate both navigation and dialog, and conduct extensive experiments analyzing the impact of different Navigator and Guide models. We highlight key challenges and publicly release the dataset, code, and evaluation framework to foster future research in embodied dialog.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。