arXiv:2511.13269cs.CV2025-11被引 14

为无人机导航设计新基准,评估视觉语言模型的空间智能

Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation

  • 构建包含13个子任务的SpatialSky-Bench基准,覆盖环境感知与场景理解
  • 在百万级数据集上训练的Sky-VLM模型表现最优,显著提升空间推理能力
  • 适合研究无人机智能、视觉语言模型或空间认知的学者与工程师

视觉语言模型(VLMs)凭借强大的视觉感知与推理能力,已被广泛应用于无人机(UAV)任务。然而,现有VLM在无人机动态环境中的空间智能能力仍缺乏系统评估,其实际导航效能存疑。为此,我们提出SpatialSky-Bench——一个专为无人机导航设计的综合性空间智能评测基准,涵盖环境感知与场景理解两大类,共13个子任务,包括边界框、颜色、距离、高度及着陆安全性分析等。对主流开源与闭源VLM的全面评估显示,在复杂无人机场景中性能普遍不理想,暴露出显著的能力差距。为应对挑战,我们构建了包含100万样本的SpatialSky-Dataset,涵盖多样化场景与多维度标注。基于此数据集,我们提出Sky-VLM,一种面向无人机多粒度、多情境空间推理的专用视觉语言模型。大量实验表明,Sky-VLM在所有基准任务中均达到当前最优水平,为开发适用于无人机场景的VLM提供了新路径。代码已开源:https://github.com/linglingxiansen/SpatialSKy。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs), leveraging their powerful visual perception and reasoning capabilities, have been widely applied in Unmanned Aerial Vehicle (UAV) tasks. However, the spatial intelligence capabilities of existing VLMs in UAV scenarios remain largely unexplored, raising concerns about their effectiveness in navigating and interpreting dynamic environments. To bridge this gap, we introduce SpatialSky-Bench, a comprehensive benchmark specifically designed to evaluate the spatial intelligence capabilities of VLMs in UAV navigation. Our benchmark comprises two categories-Environmental Perception and Scene Understanding-divided into 13 subcategories, including bounding boxes, color, distance, height, and landing safety analysis, among others. Extensive evaluations of various mainstream open-source and closed-source VLMs reveal unsatisfactory performance in complex UAV navigation scenarios, highlighting significant gaps in their spatial capabilities. To address this challenge, we developed the SpatialSky-Dataset, a comprehensive dataset containing 1M samples with diverse annotations across various scenarios. Leveraging this dataset, we introduce Sky-VLM, a specialized VLM designed for UAV spatial reasoning across multiple granularities and contexts. Extensive experimental results demonstrate that Sky-VLM achieves state-of-the-art performance across all benchmark tasks, paving the way for the development of VLMs suitable for UAV scenarios. The source code is available at https://github.com/linglingxiansen/SpatialSKy.

无人机导航视觉语言模型空间推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。