构建百万级遥感雷达图文数据集,提升AI对雷达图像的理解能力
SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding
- 构建跨城市、多分辨率的百万级SAR图像与文本配对数据集
- 在主流视觉语言模型上微调后,性能接近人类专家水平
- 适合遥感、AI多模态、地球观测领域研究者使用
合成孔径雷达(SAR)是一种关键的遥感技术,可在全天候、昼夜条件下实现地表穿透,用于精确连续的环境监测。然而,由于其复杂的成像机制与人类视觉差异大,SAR图像理解仍具挑战性。尽管视觉语言模型(VLMs)在RGB图像理解中表现优异,但其在SAR图像上的应用受限于缺乏SAR特定知识。为此,我们提出SARLANG-1M,一个面向多模态SAR图像理解的大规模基准数据集,包含超过100万张高质量的SAR图像-文本对,覆盖全球59个城市,分辨率从0.1至25米,提供细粒度语义描述(包括简短与详细标题),涵盖1,696种物体类型和16类地表覆盖类别,并支持7个应用场景、1,012种问题类型的多任务问答。在主流VLM上的大量实验表明,使用SARLANG-1M微调可显著提升模型在SAR图像理解中的表现,达到接近人类专家的水平。数据集与代码将公开发布于https://github.com/Jimmyxichen/SARLANG-1M。
原文摘要 · Abstract (English)
Synthetic Aperture Radar (SAR) is a crucial remote sensing technology, enabling all-weather, day-and-night observation with strong surface penetration for precise and continuous environmental monitoring and analysis. However, SAR image interpretation remains challenging due to its complex physical imaging mechanisms and significant visual disparities from human perception. Recently, Vision-Language Models (VLMs) have demonstrated remarkable success in RGB image understanding, offering powerful open-vocabulary interpretation and flexible language interaction. However, their application to SAR images is severely constrained by the absence of SAR-specific knowledge in their training distributions, leading to suboptimal performance. To address this limitation, we introduce SARLANG-1M, a large-scale benchmark tailored for multimodal SAR image understanding, with a primary focus on integrating SAR with textual modality. SARLANG-1M comprises more than 1 million high-quality SAR image-text pairs collected from over 59 cities worldwide. It features hierarchical resolutions (ranging from 0.1 to 25 meters), fine-grained semantic descriptions (including both concise and detailed captions), diverse remote sensing categories (1,696 object types and 16 land cover classes), and multi-task question-answering pairs spanning seven applications and 1,012 question types. Extensive experiments on mainstream VLMs demonstrate that fine-tuning with SARLANG-1M significantly enhances their performance in SAR image interpretation, reaching performance comparable to human experts. The dataset and code will be made publicly available at https://github.com/Jimmyxichen/SARLANG-1M.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。