arXiv:2410.07087cs.CVcs.RO2024-10被引 113

为无人机视觉语言导航构建真实平台与评测基准,提升导航能力。

Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology

  • 自建OpenUAV平台,支持真实飞行控制与多场景模拟
  • 构建12000条轨迹的UAV-Need-Help评测数据集
  • 提出多视图感知的分层轨迹生成方法,适合空地协同研究

开发基于语言指令和视觉信息导航至目标位置的智能体(即视觉语言导航,VLN)已引起广泛关注。现有研究多聚焦于地面智能体,而无人机(UAV)视觉语言导航仍相对薄弱。当前多数工作沿用地面设定,采用预定义离散动作空间,忽视了地面与空中环境在运动特性与任务复杂度上的本质差异。为此,本文从平台、评测基准与方法三方面提出解决方案:首先,构建支持多样化环境、真实飞行控制与算法支持的OpenUAV平台;其次,在该平台上构建包含约12,000条轨迹的目标导向型数据集,成为首个专为真实无人机VLN设计的数据集;再次,提出UAV-Need-Help辅助引导基准,提供不同层次的引导信息以应对复杂空域任务;最后,设计一种基于多模态大模型的无人机导航系统,结合多视角图像、任务描述与辅助指令,实现联合视觉文本理解与分层轨迹生成。实验表明,所提方法显著优于基线模型,但仍与人类操作者存在明显差距,凸显该任务的挑战性。

原文摘要 · Abstract (English)

Developing agents capable of navigating to a target location based on language instructions and visual information, known as vision-language navigation (VLN), has attracted widespread interest. Most research has focused on ground-based agents, while UAV-based VLN remains relatively underexplored. Recent efforts in UAV vision-language navigation predominantly adopt ground-based VLN settings, relying on predefined discrete action spaces and neglecting the inherent disparities in agent movement dynamics and the complexity of navigation tasks between ground and aerial environments. To address these disparities and challenges, we propose solutions from three perspectives: platform, benchmark, and methodology. To enable realistic UAV trajectory simulation in VLN tasks, we propose the OpenUAV platform, which features diverse environments, realistic flight control, and extensive algorithmic support. We further construct a target-oriented VLN dataset consisting of approximately 12k trajectories on this platform, serving as the first dataset specifically designed for realistic UAV VLN tasks. To tackle the challenges posed by complex aerial environments, we propose an assistant-guided UAV object search benchmark called UAV-Need-Help, which provides varying levels of guidance information to help UAVs better accomplish realistic VLN tasks. We also propose a UAV navigation LLM that, given multi-view images, task descriptions, and assistant instructions, leverages the multimodal understanding capabilities of the MLLM to jointly process visual and textual information, and performs hierarchical trajectory generation. The evaluation results of our method significantly outperform the baseline models, while there remains a considerable gap between our results and those achieved by human operators, underscoring the challenge presented by the UAV-Need-Help task.

无人机导航视觉语言导航多模态仿真平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。