用轻量级Transformer融合哨兵1/2数据,提升全球土地利用分类精度。
SpecSAR-Former: A Lightweight Transformer-based Network for Global LULC Mapping Using Integrated Sentinel-1 and Sentinel-2
- 提出双模增强与互模聚合模块,分步融合多光谱与雷达数据。
- 在动态世界+数据集上达59.58% mIoU,仅2670万参数。
- 适合遥感、地理信息、环境监测领域研究者参考。
近年来遥感领域日益关注多模态数据,得益于多样化地球观测数据的可用性。融合不同模态的互补信息显著提升了语义理解能力。然而现有全球多模态数据集常缺少合成孔径雷达(SAR)数据,而SAR在捕捉纹理与结构细节方面具有优势。为弥补此缺口,我们构建了动态世界+数据集,将对齐的SAR数据扩展至权威多光谱数据集Dynamic World。同时,为促进多光谱与SAR数据融合,提出轻量级Transformer模型SpecSAR-Former,包含双模增强模块(DMEM)和互模聚合模块(MMAM),采用分阶段融合策略挖掘两模态间交叉信息。此外,采用不平衡参数分配策略,依据模态重要性与信息密度分配参数。大量实验表明,该模型优于现有Transformer与CNN模型,在仅2670万参数下实现59.58% mIoU、79.48% OA与71.68% F1 Score。
原文摘要 · Abstract (English)
Recent approaches in remote sensing have increasingly focused on multimodal data, driven by the growing availability of diverse earth observation datasets. Integrating complementary information from different modalities has shown substantial potential in enhancing semantic understanding. However, existing global multimodal datasets often lack the inclusion of Synthetic Aperture Radar (SAR) data, which excels at capturing texture and structural details. SAR, as a complementary perspective to other modalities, facilitates the utilization of spatial information for global land use and land cover (LULC). To address this gap, we introduce the Dynamic World+ dataset, expanding the current authoritative multispectral dataset, Dynamic World, with aligned SAR data. Additionally, to facilitate the combination of multispectral and SAR data, we propose a lightweight transformer architecture termed SpecSAR-Former. It incorporates two innovative modules, Dual Modal Enhancement Module (DMEM) and Mutual Modal Aggregation Module (MMAM), designed to exploit cross-information between the two modalities in a split-fusion manner. These modules enhance the model's ability to integrate spectral and spatial information, thereby improving the overall performance of global LULC semantic segmentation. Furthermore, we adopt an imbalanced parameter allocation strategy that assigns parameters to different modalities based on their importance and information density. Extensive experiments demonstrate that our network outperforms existing transformer and CNN-based models, achieving a mean Intersection over Union (mIoU) of 59.58%, an Overall Accuracy (OA) of 79.48%, and an F1 Score of 71.68% with only 26.70M parameters. The code will be available at https://github.com/Reagan1311/LULC_segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。