提升无人机视觉语言导航的语义对齐与长期决策能力
From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation

- 引入语义增强模块,强化指令相关地标在视觉中的定位
- 动态聚合历史帧,突出高相关性关键帧生成结构化提示
- 结合拓扑感知与多奖励机制,稳定长程导航决策
无人机视觉语言导航(UAV-VLN)旨在使空中智能体根据自然语言指令,在开放三维环境中从第一人称视觉观测中完成导航。现有方法存在三大耦合问题:指令相关地标的视觉定位弱、长时序历史信息利用不足,以及在局部陷阱或重复探索中决策不稳定。为此,我们提出一个统一的语义到决策框架。首先,设计指令引导的语义增强模块,将物体级语义和相对空间线索注入当前观测状态。其次,提出相关性感知的动态时间聚合策略,重加权完整历史缓冲区,并将少数高相关帧转化为结构化地标提示输入解码器。最后,构建拓扑感知决策方法,融合局部最优认知与群体相对策略优化,在进度、目标、语义和路径合规奖励下实现稳定决策。在广泛使用的AerialVLN和OpenFly基准上的实验表明,该方法达到当前最优性能。
原文摘要 · Abstract (English)
UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environments from egocentric visual observations. Current approaches suffer from three coupled issues: weak grounding of instruction-relevant landmarks in visual observations, insufficient exploitation of long-horizon history, and unstable decisions under local traps or repeated exploration. To address these issues, we propose a unified semantic-to-decision framework. First, we present an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state. Subsequently, we develop a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frames into structured landmark prompts for the decoder. Finally, we devise a topology-aware decision method that combines local-optimum cognition with group-relative policy optimization under progress, goal, semantic, and path-compliance rewards. Experiments on the widely used AerialVLN and OpenFly benchmarks clearly demonstrate that our method achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。