arXiv:2508.15232cs.CV2025-08中稿 · ACM MM 2025被引 21

两架无人机协作,让飞行器更准地听懂指令找目标。

AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation

  • 高低空无人机分工:高空负责看全局,低空负责精准导航。
  • 在13,838条轨迹上训练,模型能适应新地图和新物体。
  • 只需交换坐标信息,效率高,适合真实户外场景部署。

空中视觉语言导航(VLN)是一项新兴任务,使无人机能根据自然语言指令和视觉线索在户外环境中导航。由于无人机路径长且机动性强,实现可靠性能常需人工干预或过于详细的指令。为此,我们提出双高度无人机协同导航任务(DuAl-VLN),由一架高空无人机负责宏观环境推理,另一架低空无人机执行精确导航。为支持训练与评估,构建了包含13,838条协同飞行轨迹的HaL-13k数据集,涵盖未见地图与未见物体验证集,系统评估模型在新环境和陌生目标下的泛化能力。我们提出AeroDuo框架,高空无人机使用多模态大语言模型(Pilot-LLM)进行目标推理,低空无人机采用轻量级多阶段策略完成导航与目标定位。二者协同仅交换少量坐标信息,保障高效性。

原文摘要 · Abstract (English)

Aerial Vision-and-Language Navigation (VLN) is an emerging task that enables Unmanned Aerial Vehicles (UAVs) to navigate outdoor environments using natural language instructions and visual cues. However, due to the extended trajectories and complex maneuverability of UAVs, achieving reliable UAV-VLN performance is challenging and often requires human intervention or overly detailed instructions. To harness the advantages of UAVs' high mobility, which could provide multi-grained perspectives, while maintaining a manageable motion space for learning, we introduce a novel task called Dual-Altitude UAV Collaborative VLN (DuAl-VLN). In this task, two UAVs operate at distinct altitudes: a high-altitude UAV responsible for broad environmental reasoning, and a low-altitude UAV tasked with precise navigation. To support the training and evaluation of the DuAl-VLN, we construct the HaL-13k, a dataset comprising 13,838 collaborative high-low UAV demonstration trajectories, each paired with target-oriented language instructions. This dataset includes both unseen maps and an unseen object validation set to systematically evaluate the model's generalization capabilities across novel environments and unfamiliar targets. To consolidate their complementary strengths, we propose a dual-UAV collaborative VLN framework, AeroDuo, where the high-altitude UAV integrates a multimodal large language model (Pilot-LLM) for target reasoning, while the low-altitude UAV employs a lightweight multi-stage policy for navigation and target grounding. The two UAVs work collaboratively and only exchange minimal coordinate information to ensure efficiency.

无人机导航视觉语言多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。