构建大规模真实航拍视觉语言导航数据集,支持更真实的无人机智能导航训练。
AirNav: A Large-Scale UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions
- 基于真实城市航拍数据,用人类与大模型协作生成自然多样指令。
- 包含13.7万条样本,新模型在未见测试集上达到51.82%成功率。
- 适合研究无人机导航、多模态大模型及真实场景迁移的团队使用。
现有无人机视觉语言导航(VLN)基准普遍缺乏真实空域场景、自然过程级指令和足够规模,难以在真实环境下系统训练与评估无人机导航智能体。为此,我们提出 extbf{AirNav},一个基于真实城市航拍数据构建的大规模基准,包含137,000个导航样本,其自然且多样的指令通过人-大模型协作管道生成,涵盖10种用户角色。我们在统一指标与开源实现下对代表性方法(从传统模型到多模态大模型)进行系统评估。进一步提出 extbf{AirVLN-R1},通过监督微调(SFT)与强化微调(RFT)训练,在测试未见分割上取得51.82%的成功率,达当前最优。物理无人机平台上的实机实验初步验证了从仿真到现实的迁移能力。数据集与代码已公开。
原文摘要 · Abstract (English)
Existing UAV vision-and-language navigation (VLN) benchmarks rarely provide realistic aerial scenes, natural process-level instructions, and sufficient scale simultaneously, making it difficult to systematically train and evaluate UAV VLN agents under realistic settings. To address this, we propose \textbf{AirNav}, a large-scale benchmark built on real urban aerial data, comprising 137K navigation samples with natural and diverse instructions generated via a human--LLM collaborative pipeline with 10 user personas. We conduct a systematic evaluation of representative approaches on AirNav, ranging from traditional models to multimodal large language models (MLLMs), under unified metrics with open-source implementations. We further propose \textbf{AirVLN-R1}, trained via supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT), achieving state-of-the-art performance with a 51.82\% success rate on the test-unseen split. Real-world experiments on a physical UAV platform provide preliminary evidence of sim-to-real transferability, and our dataset and code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。