arXiv:2608.12835cs.RO2026-08中稿 · ACM Multimedia 202…

让无人机通过当前视角预测未来空间地图,实现更精准的导航决策。

AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN

论文配图:AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN
图 1 · 摘自论文原文
  • 从多视角图像构建结构化当前地图,联合优化地图重建与未来轨迹预测。
  • 在OpenUAV和AerialVLN-S上达到新基准,相对基线提升12.3%成功率。
  • 适合需要未来感知与空间一致性建模的无人机视觉语言导航任务。

无人飞行器视觉语言导航(UAV-VLN)要求智能体根据语言指令,从稀疏的多视角观测中推断空间结构,并在复杂户外环境中执行可行的3D运动。尽管大语言模型取得进展,现有方法仍直接将视觉语言输入映射到动作,缺乏显式场景定位和未来感知的空间推理。本文提出AirForesight,一种面向UAV-VLN的当前到未来空间地图想象框架。该框架首先从多视角观测中学习结构化当前地图表示,该表示由当前地图重建和未来轨迹预测联合监督,以编码当前场景结构与未来运动意图。通过结构化因果注意力机制,当前空间知识被传递至未来地图推理,当前与未来表示聚合后预测下一3D航点。为增强空间想象与导航的相关性,引入跨空间规划一致性损失,促使预测地图轨迹方向与基于真值航点位移得到的专家动作方向保持一致。在OpenUAV和AerialVLN-S上的实验及详尽消融分析表明,所提框架性能优异,有效且稳定。

原文摘要 · Abstract (English)

Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, infer spatial structure from sparse multi-view observations, and execute feasible 3D motion in complex outdoor environments. Despite recent progress with large language models, most existing methods still map vision-language inputs directly to actions, providing limited explicit scene grounding and future-aware spatial reasoning. We propose AirForesight, a current-to-future spatial map imagination framework for UAV-VLN. AirForesight first learns a structured current-map representation from multi-view observations. This representation is jointly supervised by current-map reconstruction and future-trajectory prediction, encouraging it to encode both present scene structure and future motion intent. Under structured causal attention, the current spatial knowledge is propagated to future-map reasoning, and the resulting current and future representations are aggregated to predict the next 3D waypoint. To make spatial imagination more relevant to navigation, we introduce a cross-space planning consistency loss that encourages directional agreement between the predicted map-space trajectory and the expert action direction derived from the ground-truth waypoint displacement. Experiments on OpenUAV and AerialVLN-S, together with extensive ablations, demonstrate strong performance and support the effectiveness and stability of the proposed framework.

无人机导航空间推理视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。