构建首个统一评估无人机图像理解的基准,提出无需训练的多智能体系统提升推理准确率。
Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

- 设计多智能体系统,按任务分配视觉工具、迭代验证中间结果、动态调整搜索深度。
- 在1500个标注问题上实现77.0%准确率,超越Gemini 3 Pro 4.0%,8B模型提升8.7%。
- 针对视角混乱、尺度差异大等难题,适合无人机视觉与多模态推理研究者使用。
基于多模态大语言模型(MLLM)的无人机航拍图像理解与推理对空中智能至关重要,但面临极端尺度变化、任意相机朝向和高密度目标等挑战。现有评估分散于单一数据集与窄任务,缺乏对无人机理解与推理能力的统一衡量。为此,我们构建了UAVQA-Bench,涵盖13个公开无人机数据集的1500个人工标注问答对,覆盖6个能力维度与16项任务,支持多项选择与视觉定位两种形式。对多种开源与闭源MLLM及代理系统的系统性评估揭示三大失效模式:领域-工具错配、错误传播无控、静态推理。受此启发,我们提出UAV-MAS,一种无需训练的多智能体系统,包含:领域特定感知引擎(DSPE),用于将查询路由至适配的视觉工具;上下文感知迭代精炼模块(CAIR),通过验证中间推理抑制误差累积;难度自适应搜索机制(DAAS),根据问题难度调节搜索深度。采用32B开源MLLM的UAV-MAS在UAVQA-Bench上达到77.0%整体准确率,超过Gemini 3 Pro 4.0%;8B版本相较基线提升8.7%。
原文摘要 · Abstract (English)
Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme scale variation, arbitrary camera orientations, and high object density. Despite growing interest, existing evaluations remain fragmented across individual datasets and narrow tasks, leaving a critical gap in unified assessment of UAV understanding and reasoning capabilities. To fill this gap, we construct UAVQA-Bench, a benchmark of 1,500 human-annotated QA pairs drawn from 13 public UAV datasets, covering 6 capability dimensions and 16 tasks in both multiple-choice and visual grounding formats. Systematic evaluation of a broad range of open-source and closed-source MLLMs as well as agent-based systems on UAVQA-Bench identifies three key failure modes: domain-toolset mismatch, unchecked error propagation, and static reasoning. Motivated by these findings, we propose UAV-MAS, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine (DSPE) that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error accumulation, and a Difficulty-Aware Adaptive Search mechanism (DAAS) that adjusts search depth to question difficulty. UAV-MAS with a 32B open-source MLLM achieves 77.0% overall accuracy on UAVQA-Bench, surpassing Gemini 3 Pro by 4.0\%, while the 8B variant improves 8.7\% over its base model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。