arXiv:2512.14222cs.CVcs.RO2025-12AAAI被引 7

通过历史增强双阶段框架,提升无人机航拍视觉语言导航的精准定位能力。

History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation

  • 分粗粒度到细粒度两阶段导航,融合空间地标与历史上下文
  • 构建动态网格地图,实现视觉特征的结构化时空记忆
  • 在优化后的CityNav数据集上显著提升定位准确率,适合复杂城市导航

航拍视觉语言导航(AVLN)要求无人机在大规模城市环境中根据语言指令定位目标。成功导航需兼顾全局环境推理与局部场景理解,但现有无人机代理多采用单一粒度框架,难以平衡二者。为此,本文提出历史增强双阶段变换器(HETT)框架,通过粗粒度到细粒度的导航流程整合两者。具体而言,HETT首先融合空间地标与历史上下文预测粗略目标位置,再通过细粒度视觉分析精炼动作决策。此外,设计了一种动态网格地图,将视觉特征动态聚合为结构化空间记忆,增强整体场景感知能力。同时,对CityNav数据集标注进行人工精细化处理以提升数据质量。在优化后的CityNav数据集上实验表明,HETT取得显著性能提升,大量消融实验进一步验证了各组件的有效性。

原文摘要 · Abstract (English)

Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both global environmental reasoning and local scene comprehension, existing UAV agents typically adopt mono-granularity frameworks that struggle to balance these two aspects. To address this limitation, this work proposes a History-Enhanced Two-Stage Transformer (HETT) framework, which integrates the two aspects through a coarse-to-fine navigation pipeline. Specifically, HETT first predicts coarse-grained target positions by fusing spatial landmarks and historical context, then refines actions via fine-grained visual analysis. In addition, a historical grid map is designed to dynamically aggregate visual features into a structured spatial memory, enhancing comprehensive scene awareness. Additionally, the CityNav dataset annotations are manually refined to enhance data quality. Experiments on the refined CityNav dataset show that HETT delivers significant performance gains, while extensive ablation studies further verify the effectiveness of each component.

视觉语言导航无人机时空记忆城市导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。