构建130万对高精度遥感图像数据集,支持多任务跨模态研究。
SOMA-1M: A Large-Scale SAR-Optical Multi-resolution Alignment Dataset for Multi-Task Remote Sensing
- 设计粗到精图像匹配框架,实现雷达与光学图像像素级对齐。
- 覆盖0.5至10米分辨率,含12类地物,支持多尺度建模与训练。
- 适用于遥感多模态模型训练,推动智能解译与跨模态融合研究。
合成孔径雷达(SAR)与光学影像具有互补优势,是突破单模态限制、实现跨模态协同处理与智能解译的关键基础。然而,现有基准数据集常存在单一空间分辨率、数据规模不足、对齐精度低等问题,难以支撑多尺度基础模型的训练与泛化。为此,我们提出SOMA-1M(SAR-Optical Multi-resolution Alignment),一个包含超过130万对地理配准图像的像素级精准对齐数据集,每张图像规格为512×512像素。该数据集整合了Sentinel-1、PIESAT-1、Capella Space和Google Earth影像,实现0.5米至10米的全球多尺度覆盖,涵盖12类典型地物类别,有效保障场景多样性与复杂性。为应对多模态投影变形与海量数据配准挑战,我们设计了严格的粗到精图像匹配框架,确保像素级对齐。基于此数据集,我们建立了涵盖图像匹配、图像融合、SAR辅助云去除及跨模态翻译四大层级视觉任务的综合评估基准,涉及30余种主流算法。实验表明,在SOMA-1M上进行监督训练可显著提升各项任务性能,其中多模态遥感图像(MRSI)匹配达到当前最优水平。SOMA-1M将成为鲁棒多模态算法与遥感基础模型的重要资源。数据集将公开发布于:https://github.com/PeihaoWu/SOMA-1M。
原文摘要 · Abstract (English)
Synthetic Aperture Radar (SAR) and optical imagery provide complementary strengths that constitute the critical foundation for transcending single-modality constraints and facilitating cross-modal collaborative processing and intelligent interpretation. However, existing benchmark datasets often suffer from limitations such as single spatial resolution, insufficient data scale, and low alignment accuracy, making them inadequate for supporting the training and generalization of multi-scale foundation models. To address these challenges, we introduce SOMA-1M (SAR-Optical Multi-resolution Alignment), a pixel-level precisely aligned dataset containing over 1.3 million pairs of georeferenced images with a specification of 512 x 512 pixels. This dataset integrates imagery from Sentinel-1, PIESAT-1, Capella Space, and Google Earth, achieving global multi-scale coverage from 0.5 m to 10 m. It encompasses 12 typical land cover categories, effectively ensuring scene diversity and complexity. To address multimodal projection deformation and massive data registration, we designed a rigorous coarse-to-fine image matching framework ensuring pixel-level alignment. Based on this dataset, we established comprehensive evaluation benchmarks for four hierarchical vision tasks, including image matching, image fusion, SAR-assisted cloud removal, and cross-modal translation, involving over 30 mainstream algorithms. Experimental results demonstrate that supervised training on SOMA-1M significantly enhances performance across all tasks. Notably, multimodal remote sensing image (MRSI) matching performance achieves current state-of-the-art (SOTA) levels. SOMA-1M serves as a foundational resource for robust multimodal algorithms and remote sensing foundation models. The dataset will be released publicly at: https://github.com/PeihaoWu/SOMA-1M.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。