arXiv:2603.08007cs.CVcs.AI2026-03被引 2

通过视觉空间推理提升无人机视觉语言导航能力

ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation

  • 设计三阶段协同架构,用视觉提示引导模型直接在图像上推理
  • 在CityNav基准上成功率比现有最佳方法提高70.3%
  • 无需额外训练,适合构建高效无人机导航系统

现有的空中视觉语言导航(Aerial VLN)方法主要采用检测-规划流水线,将开放词汇检测结果转换为离散的文本场景图。这类方法存在空间推理能力不足和固有的语言歧义问题。为此,我们提出一种视觉-空间推理(ViSA)增强框架用于空中VLN。具体而言,设计了一个三阶段协同架构,利用结构化视觉提示,使视觉语言模型(VLMs)能够在不依赖额外训练或复杂中间表示的情况下,直接在图像平面上进行推理。在CityNav基准上的全面评估表明,该ViSA增强的VLN相比完全训练的最先进方法,成功率提升了70.3%,凸显其作为空中VLN系统骨干的巨大潜力。

原文摘要 · Abstract (English)

Existing aerial Vision-Language Navigation (VLN) methods predominantly adopt a detection-and-planning pipeline, which converts open-vocabulary detections into discrete textual scene graphs. These approaches are plagued by inadequate spatial reasoning capabilities and inherent linguistic ambiguities. To address these bottlenecks, we propose a Visual-Spatial Reasoning (ViSA) enhanced framework for aerial VLN. Specifically, a triple-phase collaborative architecture is designed to leverage structured visual prompting, enabling Vision-Language Models (VLMs) to perform direct reasoning on image planes without the need for additional training or complex intermediate representations. Comprehensive evaluations on the CityNav benchmark demonstrate that the ViSA-enhanced VLN achieves a 70.3\% improvement in success rate compared to the fully trained state-of-the-art (SOTA) method, elucidating its great potential as a backbone for aerial VLN systems.

视觉导航空间推理无人机多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。