arXiv:2603.21046cs.CVcs.AI2026-03

通过几何先验引导视觉重参数化,提升无人机视觉语言导航的三维空间推理能力。

SpatialFly: Implicit 3D Prior-Guided Visual Reparameterization for Continuous UAV Vision-and-Language Navigation

  • 利用几何先验注入2D语义令牌,提供场景级结构指导。
  • 在未见环境中路径误差降低4.03米,成功率提升1.27%。
  • 适合需要精准三维导航的无人机自主探索任务。

无人机在自主探索、灾害响应和基础设施检查等应用中发挥重要作用,但在复杂三维环境中的视觉语言导航(UAV VLN)仍具挑战性。核心难点在于二维视觉感知与三维轨迹决策空间之间的结构表征不匹配,限制了空间推理能力。为此,我们提出SpatialFly,一种面向无人机视觉语言导航的几何引导空间表征框架。该框架基于RGB观测,无需显式三维重建,引入几何引导的二维自适应表征机制:几何先验注入模块将全局结构线索注入2D语义令牌,提供场景级几何指导;几何感知重参数化模块则通过几何条件化的跨模态注意力与门控残差融合,自适应地重参数化视觉令牌。实验结果表明,SpatialFly在可见与不可见环境上均持续优于现有最优基线,在未见环境的Full划分上,路径误差(NE)降低4.03米,成功率(SR)提升1.27%。轨迹级分析显示,SpatialFly生成的路径具有更好的路径对齐性与更平滑稳定的运动表现。

原文摘要 · Abstract (English)

UAVs play an important role in applications such as autonomous exploration, disaster response, and infrastructure inspection. However, UAV VLN in complex 3D environments remains challenging. A key difficulty is the structural representation mismatch between 2D visual perception and the 3D trajectory decision space, which limits spatial reasoning. To this end, we propose SpatialFly, a geometry-guided spatial representation framework for UAV VLN. Operating on RGB observations without explicit 3D reconstruction, SpatialFly introduces a geometry-guided 2D adaptive representation mechanism. Specifically, the geometric prior injection module injects global structural cues into 2D semantic tokens to provide scene-level geometric guidance. The geometry-aware reparameterization module then uses geometry-conditioned cross-modal attention and gated residual fusion to adaptively reparameterize the visual tokens. Experimental results show that SpatialFly consistently outperforms state-of-the-art UAV VLN baselines across both seen and unseen environments, reducing NE by 4.03m and improving SR by 1.27% over the strongest baseline on the unseen Full split. Additional trajectory-level analysis shows that SpatialFly produces trajectories with better path alignment and smoother, more stable motion.

无人机导航视觉语言几何先验空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。