评测大模型在低空无人机场景下的感知、认知与规划能力
MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
- 构建涵盖19个子任务的综合评测基准,覆盖真实无人机数据
- 5.7K人工标注问题揭示当前模型在复杂场景中表现不佳
- 发现空间偏差和多视角理解是关键瓶颈,适合无人机研发者参考
尽管多模态大语言模型(MLLMs)在多个领域展现出卓越的通用智能,但其在以无人机(UAVs)为主导的低空应用中的潜力仍基本未被探索。现有MLLM基准很少涉及低空场景的独特挑战,而无人机相关评估通常局限于定位或导航等特定任务,缺乏对MLLM通用智能的统一评价。为填补这一空白,我们提出MM-UAVBench,一个系统性评估MLLM在低空无人机场景中感知、认知与规划三方面能力的综合性基准。该基准包含19个子任务,超过5.7K条人工标注问题,均源自公开数据集中的真实无人机数据。对16个开源及专有MLLM的广泛实验表明,当前模型难以适应低空场景复杂的视觉与认知需求。进一步分析揭示了空间偏差和多视角理解等关键瓶颈,制约了MLLM在无人机场景中的有效部署。我们希望MM-UAVBench能推动面向真实世界无人机智能的鲁棒可靠模型研究。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) have exhibited remarkable general intelligence across diverse domains, their potential in low-altitude applications dominated by Unmanned Aerial Vehicles (UAVs) remains largely underexplored. Existing MLLM benchmarks rarely cover the unique challenges of low-altitude scenarios, while UAV-related evaluations mainly focus on specific tasks such as localization or navigation, without a unified evaluation of MLLMs'general intelligence. To bridge this gap, we present MM-UAVBench, a comprehensive benchmark that systematically evaluates MLLMs across three core capability dimensions-perception, cognition, and planning-in low-altitude UAV scenarios. MM-UAVBench comprises 19 sub-tasks with over 5.7K manually annotated questions, all derived from real-world UAV data collected from public datasets. Extensive experiments on 16 open-source and proprietary MLLMs reveal that current models struggle to adapt to the complex visual and cognitive demands of low-altitude scenarios. Our analyses further uncover critical bottlenecks such as spatial bias and multi-view understanding that hinder the effective deployment of MLLMs in UAV scenarios. We hope MM-UAVBench will foster future research on robust and reliable MLLMs for real-world UAV intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。