arXiv:2607.21400cs.ROcs.AI2026-07

让无人机仅靠视觉线索完成长距离导航,摆脱对语言指令的依赖。

VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method

论文配图:VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method
图 1 · 摘自论文原文
  • 构建纯视觉线索导航范式,将路径信息从语言指令转移到场景中可感知的线索。
  • 在7210个任务中测试,最困难场景成功率达1.8%,揭示长期证据融合挑战。
  • 适合研究视觉导航、自主飞行与多模态学习的学者参考。

视觉-语言导航(VLN)使具身智能体能够遵循自然语言指令。然而,路线级指令通常包含空间先验(如朝向、距离、布局),这些在开放、无GPS环境部署时无法通过机载传感器直接获取。因此,现有基准性能同时反映视觉导航能力与任务描述提供的结构化信息利用程度。为此,我们提出纯视觉长时程导航(VoLN),将路线相关信息从外部语言指令和全局引导转移至局部可观测的场景线索。在VoLN中,目标视图定义目的地,路径相关信息仅通过智能体需在线检测、解读并选择的局部场景线索获得。我们为航拍导航构建了VoLN-UAV基准,包含7210个任务,涵盖长时程定向飞行、连续3D运动、大幅视角变化及上下文依赖的信标选择。此外,我们提供初始基线模型VoLN-MLLM,其将自监督视觉特征与结构化语义空间对齐,并基于观测历史、目标视图、检索到的视觉-语义标记和本体感知预测短时程航点段。在五个环境的Test-Unseen划分上,该模型在Easy、Normal、Hard任务中的成功率分别为7.4%、4.5%和1.8%。这些结果为VoLN提供了初步评估,揭示了长时程证据整合、跨视图目标匹配与闭环稳定性等方面的显著挑战。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encode spatial priors, such as orientation, distance, and layout, that are not explicitly available from onboard sensing at deployment in open, GPS-denied environments. Benchmark performance under such interfaces therefore jointly reflects visual navigation ability and the use of route structure explicitly supplied by the task description. As a complementary formulation, we propose Vision-Only Long-Horizon Navigation (VoLN), which shifts route-relevant information from externally supplied instructions and global guidance to locally observable in-scene cues. In VoLN, goal views specify the destination, while route-relevant information is available only through locally observable in-scene cues that the agent must detect, interpret, and select online. We instantiate VoLN for aerial navigation through VoLN-UAV, a 7,210-episode benchmark that combines long-horizon goal-directed flight, continuous 3D motion, large viewpoint changes, and context-dependent beacon selection. We further provide VoLN-MLLM as an initial reference baseline. It aligns self-supervised visual features with a structured semantic space and predicts short-horizon waypoint segments from observation history, goal views, retrieved visual--semantic tokens, and proprioception. On the five-environment Test-Unseen split, it obtains success rates of 7.4%, 4.5%, and 1.8% on Easy, Normal, and Hard episodes, respectively. These results provide an initial evaluation of VoLN and reveal substantial remaining challenges in long-horizon evidence integration, cross-view goal matching, and closed-loop stability. Project page: https://admire-ljb.github.io/VoLN-UAV/

视觉导航无人机长程规划自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。