arXiv:2604.02020cs.CV2026-04被引 2

首个评估无人机与卫星跨视图空间智能的基准,破解视觉语言模型在空地协同中的定位难题。

Are VLMs Lost Between Sky and Space? LinkS$^2$Bench for UAV-Satellite Dynamic Cross-View Spatial Intelligence

  • 构建空地动态跨视图数据集,连接超200km²卫星影像与1022分钟无人机视频。
  • 18个主流模型在17.9k问答对上表现远低于人类,跨视图动态对齐是核心瓶颈。
  • 提出显式对齐适配器,显著提升模型在空地联动任务中的推理能力,适合遥感与AI融合研究者。

无人机与卫星的协同空间智能对应急响应和安全行动至关重要,因其兼具宏观全局覆盖与实时动态局部感知。然而,现有视觉语言模型(VLMs)在处理这种复杂跨视图交互方面的能力仍不明晰,主要因现有基准仅限于孤立的无人机视频或静态卫星图像,无法评估动态本地到全局的空间映射能力。为此,我们提出LinkS²Bench,首个专为评估VLMs广域动态跨视图空间智能设计的基准。该数据集将1,022分钟动态无人机影像与覆盖超200 km²的高分辨率卫星影像相链接。通过基于大语言模型的辅助流程与严格人工标注,构建了包含17.9k高质量问答对的评测集,涵盖感知、定位、关系与推理四大维度的12项细粒度任务。对18个代表性VLMs的评估显示,其性能显著落后于人类基线,准确的跨视图动态对齐被识别为关键瓶颈。为此,我们设计了跨视图对齐适配器,证明显式对齐能有效提升模型表现。此外,微调实验表明LinkS²Bench在推动VLM适应复杂空间推理方面的潜力。

原文摘要 · Abstract (English)

Synergistic spatial intelligence between UAVs and satellites is indispensable for emergency response and security operations, as it uniquely integrates macro-scale global coverage with dynamic, real-time local perception. However, the capacity of Vision-Language Models (VLMs) to master this complex interplay remains largely unexplored. This gap persists primarily because existing benchmarks are confined to isolated Unmanned Aerial Vehicle (UAV) videos or static satellite imagery, failing to evaluate the dynamic local-to-global spatial mapping essential for comprehensive cross-view reasoning. To bridge this gap, we introduce LinkS$^2$Bench, the first comprehensive benchmark designed to evaluate VLMs' wide-area, dynamic cross-view spatial intelligence. LinkS$^2$Bench links 1,022 minutes of dynamic UAV footage with high-resolution satellite imagery covering over 200 km$^2$. Through an LMM-assisted pipeline and rigorous human annotation, we constructed 17.9k high-quality question-answer pairs comprising 12 fine-grained tasks across four dimensions: perception, localization, relation, and reasoning. Evaluations of 18 representative VLMs reveal a substantial gap compared to human baselines, identifying accurate cross-view dynamic alignment as the critical bottleneck. To alleviate this, we design a Cross-View Alignment Adapter, demonstrating that explicit alignment significantly improves model performance. Furthermore, fine-tuning experiments underscore the potential of LinkS$^2$Bench in advancing VLM adaptation for complex spatial reasoning.

跨视图推理无人机遥感视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。