构建首个基于地籍矢量数据的遥感理解数据集,提升模型对地理细节的精准识别能力。
GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data
- 基于真实地籍数据构建380万标注对象的高分辨率图像数据集
- 在7类空间推理任务中验证模型表现,零样本下现有模型性能不足
- 仅靠高质量标注即可让通用模型实现精细空间定位,无需复杂结构改动
精确的空间理解对于地球观测至关重要,可将原始航拍影像转化为城市规划、环境监测与灾害管理等关键应用的可用信息。然而,多模态大语言模型在遥感领域存在细粒度空间理解能力不足的问题,主要源于依赖有限或重构的旧数据集。为此,我们提出一个大规模、基于可验证地籍矢量数据的数据集,包含51万张高分辨率图像中的380万个标注对象,涵盖135个细粒度语义类别。我们通过覆盖七类空间推理任务的指令微调基准进行验证,采用标准LLaVA架构建立稳健基线。实验表明,当前专用于遥感及商用模型(如Gemini)在零样本设置下表现不佳,而高保真监督能有效弥补这一差距,使标准架构在不改变结构的前提下掌握细粒度空间定位能力。
原文摘要 · Abstract (English)
Precise spatial understanding in Earth Observation is essential for translating raw aerial imagery into actionable insights for critical applications like urban planning, environmental monitoring and disaster management. However, Multimodal Large Language Models exhibit critical deficiencies in fine-grained spatial understanding within Remote Sensing, primarily due to a reliance on limited or repurposed legacy datasets. To bridge this gap, we introduce a large-scale dataset grounded in verifiable cadastral vector data, comprising 3.8 million annotated objects across 510k high-resolution images with 135 granular semantic categories. We validate this resource through a comprehensive instruction-tuning benchmark spanning seven spatial reasoning tasks. Our evaluation establishes a robust baseline using a standard LLaVA architecture. We show that while current RS-specialized and commercial models (e.g., Gemini) struggle in zero-shot settings, high-fidelity supervision effectively bridges this gap, enabling standard architectures to master fine-grained spatial grounding without complex architectural modifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。