arXiv:2608.11901cs.RO2026-08

首个支持连续动作与动态元素的室外视觉语言导航数据集

DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements

论文配图:DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements
图 1 · 摘自论文原文
  • 构建连续动作与动态元素并存的室外导航环境
  • 6933条轨迹显示任务成功率下降超10%,挑战显著
  • 适合研究真实世界智能体导航与仿真到现实的迁移

视觉-语言导航(VLN)正从室内向室外环境扩展。然而,现有室外VLN数据集仍依赖固定的离散拓扑图构建,无法反映真实世界中快速变化的环境,阻碍了导航智能体的仿真到现实迁移。为解决这一问题,我们提出DaViNCi(Dynamic Vision-and-Language Navigation in Continuous Environment),首个同时引入连续动作和动态元素的室外VLN数据集。智能体在户外环境中以连续动作移动,并需应对不可预测的动态元素。数据集包含六个不同地图,共计6,933条轨迹。通过全面对比实验发现,相较于以往数据集,DaViNCi上的成功率在离散环境下下降超过10%,在连续设置下降幅更大,凸显其挑战性。此外,我们揭示了动作粒度与动态元素的影响。这些结果证明了DaViNCi在推动室外VLN向更真实环境发展的实用价值。官网:https://xzh0312.github.io/DaViNCi/

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) has progressively expanded from indoor to outdoor environments. However, existing outdoor VLN datasets still rely on fixed discrete topological graphs for construction. It fails to align with the rapidly changing real-world outdoor environments and impedes the sim-to-real transfer of VLN agents. To address this limitation, we propose DaViNCi (\textbf{D}yn\textbf{a}mic \textbf{Vi}sion-and-Language \textbf{N}avigation in \textbf{C}ont\textbf{i}nuous Environment), the first outdoor VLN dataset that simultaneously introduces both continuous and dynamic factors. The agent not only moves in the outdoor environment using continuous actions but is also required to handle unpredictable dynamic elements. The dataset encompasses six distinct maps with a total of 6,933 trajectories. Through comprehensive comparative experiments, we find that the success rate on DaViNCi decreased by more than 10\% in discrete environments compared to previous datasets. And there is an even greater decline in continuous settings, demonstrating the challenge of DaViNCi. Furthermore, we clarify the impact of action granularity and dynamic elements. These results demonstrate the practical value of DaViNCi in advancing outdoor VLN toward more realistic environments. The website is https://xzh0312.github.io/DaViNCi/.

视觉语言导航连续动作动态环境室外场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。