arXiv:2608.01802cs.AIcs.RO2026-08

用博弈论设计双无人机协同导航,提升复杂城市环境下的目标寻访成功率。

CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning

论文配图:CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning
图 1 · 摘自论文原文
  • 构建高低空无人机的领导者-追随者博弈框架,实现分工协作。
  • 在AerialVLN数据集上成功率最高提升30.8点,跨场景迁移提升9.0点。
  • 适合做无人机自主导航、多智能体协同任务的研究与应用。

面向灾后救援、基础设施巡检和安全巡逻等任务,基于视觉与语言的目标导向空中导航正受到广泛关注。在此任务中,无人机需仅凭目标外观与环境描述定位目标,要求兼具全局探索与精准定位能力,且避免碰撞,二者难以由单一智能体兼顾。现有方法多将地面导航范式迁移至低空无人机,并依赖外部辅助提升探索效率;近期尝试部署两架无人机于不同高度,但仍依赖特权信息,且独立训练,无法实现相互适应。本文提出CoNav-UAV,将任务建模为高层领导者与低层追随者之间的斯塔克尔伯格博弈,系统仅使用机载视觉与语言输入。为求解该博弈,引入迭代斯塔克尔伯格学习:领导者通过基于记忆的上下文学习优化高层视觉-语言推理,追随者则通过类似DAgger的专家蒸馏更新精确运动控制,交替优化使双方趋向均衡。在AerialVLN基准三个高保真城市场景上,CoNav-UAV持续优于单/双智能体基线,学习场景成功率最高提升30.8点,跨场景迁移提升9.0点,且仅需约1/3的适配数据。分析验证了双智能体更新的互补性,并揭示不同视觉语言模型骨干下的差异性学习动态。

原文摘要 · Abstract (English)

Target-oriented vision-and-language navigation (VLN) on aerial platforms is attracting growing attention for missions such as disaster rescue, infrastructure inspection, and security patrol. In this task, an unmanned aerial vehicle (UAV) needs to locate targets given only a concise description of their appearance and surroundings. This requires global exploration and grounding as well as collision-free close-range approach, two interleaved processes difficult to reconcile within a single agent. Most existing methods transfer the ground VLN paradigm to a low-altitude UAV and compensate for its inefficient exploration with external assistance. A recent attempt deploys two UAVs at complementary altitudes yet still relies on privileged information and trains its two agents independently, precluding any mutual adaptation essential for cooperation. Here we propose CoNav-UAV, which explicitly models the task as a Stackelberg game between a high-altitude leader and a low-altitude follower, with the system operating on onboard visual and linguistic inputs alone. To solve this game, we introduce Iterative Stackelberg Learning. The leader's high-level vision-language reasoning is refined via memory-based in-context learning, while the follower's precise motion control is updated via DAgger-style expert distillation. The alternation drives both agents toward a Stackelberg equilibrium. CoNav-UAV consistently outperforms single- and dual-agent baselines across three high-fidelity urban scenes from the AerialVLN benchmark. Success rate improves by up to 30.8 points on the learning scene, and 9.0 points under cross-scene transfer while using about 3x less adaptation data. Further analyses validate the complementary gains of the leader and follower updates and reveal robust gains yet distinct learning dynamics across VLM backbones.

无人机导航多智能体博弈学习视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。