arXiv:2606.27876cs.CVcs.AI2026-06

构建首个低空无人机空间智能评测基准,覆盖多视角协同与动态理解。

SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion

论文配图:SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion
图 1 · 摘自论文原文
  • 设计统一视觉-问答框架,支持七种输入和九种答案格式。
  • 包含4331个实例,覆盖14类细粒度任务,涵盖空间关系与运动理解。
  • 揭示当前模型在跨视图关联与几何推理上的显著短板,适合算法评估。

空间智能对低空无人机感知、协作与导航至关重要。现有基准多聚焦图像级识别、单视图理解或狭义答案格式,难以充分评估三维空间推理、多视角协作、场景动态变化及多样化任务形式。为此,我们提出SpatialUAV——一个真实低空无人机评测基准,包含4,331个精心筛选的实例,覆盖14种细粒度任务类型,涵盖语义区分、空间关系、空中-空中协作、空中-地面协作以及运动理解。所有样本统一为视觉输入-问题-答案结构,支持七种输入配置和九种答案格式,包括选项标签、区域标识、几何数值、跨视图对应及自由文本运动描述。数据构建采用检测器辅助区域、深度监督、元数据规则、大量人工标注、盲筛及多轮人工验证,并配备任务特异性指标以实现可靠评估。在三类代表性视觉语言模型上评估发现,当前模型距离人类水平仍有显著差距,尤其在跨视图关联、结构化定位、几何推理与时间视角理解方面存在明显瓶颈。该结果为提升低空无人机空间智能提供了实证指导。代码与数据见https://github.com/Hyu-Zhang/SpatialUAV。

原文摘要 · Abstract (English)

Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation. However, existing UAV benchmarks often emphasize image-level recognition, single-view understanding, or narrow answer formats, leaving 3D spatial inference, multi-view collaboration, scene dynamics, and diverse task formulations insufficiently evaluated. To address these gaps, we introduce SpatialUAV, a real low-altitude UAV benchmark comprising 4,331 curated instances across 14 fine-grained task types, covering semantic discrimination, spatial relation, aerial--aerial collaboration, aerial--ground collaboration, and motion understanding. SpatialUAV organizes all samples into a unified visual-input--question--answer schema, while supporting seven input configurations and nine answer formats, including option labels, region identifiers, geometric values, cross-view correspondences, and free-form motion descriptions. To ensure reliable and grounded evaluation, our data construction pipeline integrates detector-assisted regions, depth supervision, metadata-derived rules, extensive manual annotation, blind filtering, and multi-turn human validation, together with task-specific metrics for heterogeneous outputs. Evaluating representative vision-language models across three categories, we show that current models remain far from human-level performance, with pronounced bottlenecks in cross-view association, structured grounding, geometric reasoning, and temporal viewpoint understanding. These results offer empirical guidance for advancing low-altitude UAV spatial intelligence. Code and data are available at https://github.com/Hyu-Zhang/SpatialUAV.

无人机空间智能多视角评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。