arXiv:2603.14336cs.CV2026-03被引 2

为低空无人机视觉语言理解打造了新基准和训练数据集

UAVBench and UAVIT-1M: Benchmarking and Enhancing MLLMs for Low-Altitude UAV Vision-Language Understanding

  • 构建了涵盖10项任务的UAVBench基准和124万条指令的UAVIT-1M数据集
  • 实测发现开源模型在低空视觉理解上显著落后于闭源模型
  • 用UAVIT-1M微调后,开源模型性能大幅提升,接近实用水平

多模态大语言模型在自然图像和卫星遥感图像上已取得显著进展,但对低空无人机场景的理解仍具挑战。现有数据集多聚焦少数特定任务,难以全面评估模型在真实低空无人机应用中的表现。为此,我们提出UAVBench综合基准和UAVIT-1M大规模指令微调数据集,用于评测与提升MLLMs在低空视觉语言任务中的能力。UAVBench包含43个测试单元、96.6万条高质量样本,覆盖图像级与区域级共10项任务;UAVIT-1M包含约124万条多样化指令,覆盖78.9万张多场景图像、约2000种空间分辨率,涉及11项任务。两者均采用纯真实世界图像,涵盖丰富天气条件,并经人工验证保证质量。对11个前沿MLLMs的深入分析显示,开源模型在低空视觉内容对话生成上准确率不足,显著落后于闭源模型。大量实验表明,在UAVIT-1M上微调可有效缩小此差距。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have made significant strides in natural images and satellite remote sensing images. However, understanding low-altitude drone scenarios remains a challenge. Existing datasets primarily focus on a few specific low-altitude visual tasks, which cannot fully assess the ability of MLLMs in real-world low-altitude UAV applications. Therefore, we introduce UAVBench, a comprehensive benchmark, and UAVIT-1M, a large-scale instruction tuning dataset, designed to evaluate and improve MLLMs' abilities in low-altitude vision-language tasks. UAVBench comprises 43 test units and 966k high-quality data samples across 10 tasks at the image-level and region-level. UAVIT-1M consists of approximately 1.24 million diverse instructions, covering 789k multi-scene images and about 2,000 types of spatial resolutions with 11 distinct tasks. UAVBench and UAVIT-1M feature pure real-world visual images and rich weather conditions, and involve manual verification to ensure high quality. Our in-depth analysis of 11 state-of-the-art MLLMs using UAVBench reveals that open-source MLLMs cannot generate accurate conversations about low-altitude visual content, lagging behind closed-source MLLMs. Extensive experiments demonstrate that fine-tuning open-source MLLMs on UAVIT-1M significantly addresses this gap. Our contributions pave the way for bridging the gap between current MLLMs and low-altitude UAV real-world application demands. (Project page: https://UAVBench.github.io/)

无人机视觉多模态模型基准测试指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。