用多模态大模型分析卫星图像中的工地活动演变过程。
Geospatial-Temporal Sensemaking of Remote Sensing Activity Detections with Multimodal Large Language Model
- 将卫星图像与标注转化为自然语言问答对,支持时空推理
- 构建包含230万组时序对比题目的大规模数据集
- 适合关注遥感智能分析与大模型融合的 researchers
本文提出 SMART-HC-VQA,一个基于 Sentinel-2 的视觉问答数据集,源自 IARPA SMART Heavy Construction 数据集,旨在支持人类活动的时空分析。该数据集将工地标注、施工类型标签、时间阶段标签、地理元数据及观测关系转换为自然语言问答三元组。此方法将原数据集重构为一个随时间扩展的自动目标识别与视觉问答挑战,以固定地理区域为目标,其属性与活动状态在稀疏卫星观测中动态变化。当前数据集包含 21,837 个可访问的 Sentinel-2 图像块、65,511 个单图 VQA 示例,以及约 230 万组通过新型图像对组合增强生成的双图时序对比示例。文中详细描述了 Sentinel-2 影像获取与处理流程、大卫星瓦片分割为以工地为中心的图像、追踪至 SMART-HC 标注的可追溯性,以及站点大小、观测次数、时间覆盖度、施工类型和阶段标签的分布分析。此外,还实现了一个基于 LLaVA-NeXT Mistral-7B 的多图像多模态大模型训练框架,可接受多个带时间戳的图像输入,并在基于元数据的 VQA 示例上进行训练。本工作为语言引导的遥感活动理解提供了可复现的基础,不仅旨在检测变化,更致力于推断正在进行的过程、其演进轨迹及潜在未来发展趋势。
原文摘要 · Abstract (English)
We introduce SMART-HC-VQA, a Sentinel-2-based visual question answering dataset derived from the IARPA SMART Heavy Construction dataset, designed for spatiotemporal analysis of human activity. The dataset transforms construction-site annotations, construction-type labels, temporal-phase labels, geographic metadata, and observation relationships into natural language question-answer triplets. This approach redefines the existing dataset as a temporally extended automatic target recognition and visual question answering (VQA) challenge, considering a fixed geospatial site as a target whose attributes and activity states evolve across sparse satellite observations. Currently, SMART-HC-VQA comprises 21,837 accessible Sentinel-2 image chips, 65,511 single-image VQA examples, and approximately 2.3 million two-image temporal comparison examples generated via our novel Image-Pairwise Combinatorial Augmentation. We detail the workflow for retrieving and processing Sentinel-2 imagery, segmenting large satellite tiles into site-centered images, maintaining traceability to SMART-HC annotations, and analyzing the distributions of site size, observation count, temporal coverage, construction type, and phase labels. Additionally, we describe an implemented multi-image MLLM training framework based on LLaVA-NeXT Mistral-7B, adapted to accept multiple dated image inputs and train on metadata-derived VQA examples. This work offers a reproducible foundation for understanding language-guided remote sensing activities, aiming not only to detect change but also to reason about the ongoing processes, their progression, and potential future developments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。