融合高分辨率影像与时间序列数据,提升遥感图像语义分割精度。
Deep Multimodal Fusion for Semantic Segmentation of Remote Sensing Earth Observation Data
- 采用双分支网络分别处理航拍影像和卫星时序数据。
- 在FLAIR数据集上达到当前最优分割效果。
- 适合遥感土地覆盖分析、城市规划等应用研究者。
遥感图像的精准语义分割对土地覆盖制图、城市规划和环境监测等地球观测应用至关重要。然而,单一数据源存在局限:超高清航空影像虽具丰富空间细节,却无法捕捉地表变化的时间信息;而卫星时序图像(SITS)能反映植被季节性等动态变化,但空间分辨率较低,难以区分细粒度目标。本文提出一种晚期融合深度学习模型(LF-DLM),融合航空影像与哨兵-2卫星时序图像的优势。模型包含两个独立分支:一个分支使用基于MaxViT骨干的UNetFormer整合航空影像中的纹理细节;另一个分支采用带时间注意力编码器的U-Net(U-TAE)捕捉SITS中的复杂时空动态。该方法在大型多源光学影像基准数据集FLAIR上取得当前最优性能,验证了多模态融合对提升遥感语义分割准确性与鲁棒性的关键作用。
原文摘要 · Abstract (English)
Accurate semantic segmentation of remote sensing imagery is critical for various Earth observation applications, such as land cover mapping, urban planning, and environmental monitoring. However, individual data sources often present limitations for this task. Very High Resolution (VHR) aerial imagery provides rich spatial details but cannot capture temporal information about land cover changes. Conversely, Satellite Image Time Series (SITS) capture temporal dynamics, such as seasonal variations in vegetation, but with limited spatial resolution, making it difficult to distinguish fine-scale objects. This paper proposes a late fusion deep learning model (LF-DLM) for semantic segmentation that leverages the complementary strengths of both VHR aerial imagery and SITS. The proposed model consists of two independent deep learning branches. One branch integrates detailed textures from aerial imagery captured by UNetFormer with a Multi-Axis Vision Transformer (MaxViT) backbone. The other branch captures complex spatio-temporal dynamics from the Sentinel-2 satellite image time series using a U-Net with Temporal Attention Encoder (U-TAE). This approach leads to state-of-the-art results on the FLAIR dataset, a large-scale benchmark for land cover segmentation using multi-source optical imagery. The findings highlight the importance of multi-modality fusion in improving the accuracy and robustness of semantic segmentation in remote sensing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。