arXiv:2604.08883cs.ROcs.AI2026-04被引 1

融合模仿与强化学习,提升城市空中导航的精度与鲁棒性。

HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation

  • 分层决策机制协同宏观路径规划与微观动作控制
  • 在CityNav上各场景层级任务均达最优表现
  • 适合城市巡检、物流配送等复杂环境应用

受通用视觉-语言导航(VLN)任务启发,空中视觉-语言导航因在物流配送和城市巡检中的重要实用价值而受到广泛关注。然而,现有方法在复杂城市环境中面临泛化能力不足、长距离路径规划性能不佳以及空间连续性理解欠缺等问题。为此,本文提出HTNav,一种融合模仿学习(IL)与强化学习(RL)的混合式导航框架。该框架采用分阶段训练机制,在保证基础导航策略稳定性的同时增强环境探索能力。通过引入分层决策机制,实现宏观路径规划与细粒度动作控制的协同交互。此外,设计了地图表征学习模块,深化对开放域空间连续性的理解。在CityNav基准测试中,本方法在所有场景层级和任务难度下均取得当前最优性能。实验表明,该框架显著提升了复杂城市环境下的导航精度与鲁棒性。

原文摘要 · Abstract (English)

Inspired by the general Vision-and-Language Navigation (VLN) task, aerial VLN has attracted widespread attention, owing to its significant practical value in applications such as logistics delivery and urban inspection. However, existing methods face several challenges in complex urban environments, including insufficient generalization to unseen scenes, suboptimal performance in long-range path planning, and inadequate understanding of spatial continuity. To address these challenges, we propose HTNav, a new collaborative navigation framework that integrates Imitation Learning (IL) and Reinforcement Learning (RL) within a hybrid IL-RL framework. This framework adopts a staged training mechanism to ensure the stability of the basic navigation strategy while enhancing its environmental exploration capability. By integrating a tiered decision-making mechanism, it achieves collaborative interaction between macro-level path planning and fine-grained action control. Furthermore, a map representation learning module is introduced to deepen its understanding of spatial continuity in open domains. On the CityNav benchmark, our method achieves state-of-the-art performance across all scene levels and task difficulties. Experimental results demonstrate that this framework significantly improves navigation precision and robustness in complex urban environments.

视觉语言导航城市巡检强化学习分层决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。