融合多源遥感数据,用新框架提升图像语义分割精度。
A Unified Framework with Multimodal Fine-tuning for Remote Sensing Semantic Segmentation
- 设计可扩展的多模态微调网络,适配不同模型优化策略。
- 在3个基准数据集上显著超越现有方法,最高提升达12.3%。
- 适合遥感分析、地理信息研究者参考,尤其关注多源数据融合。
多模态遥感数据来自多种传感器,能全面呈现地表信息。利用多模态融合技术,语义分割可实现比单模态更精细准确的地理场景分析。基于视觉基础模型(如Segment Anything Model, SAM),本文提出统一框架,包含新型多模态微调网络(MFNet),支持适配器(Adapter)和低秩适应(LoRA)等微调机制。该框架在保留SAM通用知识的同时,有效融合多模态数据。此外,引入金字塔式深度融合模块(DFM),整合多尺度高层地理特征,增强解码前的表征能力。本工作还首次验证了SAM对数字高程模型(DSM)数据的强大泛化能力。在ISPRS Vaihingen、ISPRS Potsdam和MMHunan三个基准数据集上的实验表明,所提方法显著优于现有方法,性能达到新高度,为未来研究提供通用基础。代码已开源:https://github.com/sstary/SSRS。
原文摘要 · Abstract (English)
Multimodal remote sensing data, acquired from diverse sensors, offer a comprehensive and integrated perspective of the Earth's surface. Leveraging multimodal fusion techniques, semantic segmentation enables detailed and accurate analysis of geographic scenes, surpassing single-modality approaches. Building on advancements in vision foundation models, particularly the Segment Anything Model (SAM), this study proposes a unified framework incorporating a novel Multimodal Fine-tuning Network (MFNet) for remote sensing semantic segmentation. The proposed framework is designed to seamlessly integrate with various fine-tuning mechanisms, demonstrated through the inclusion of Adapter and Low-Rank Adaptation (LoRA) as representative examples. This extensibility ensures the framework's adaptability to other emerging fine-tuning strategies, allowing models to retain SAM's general knowledge while effectively leveraging multimodal data. Additionally, a pyramid-based Deep Fusion Module (DFM) is introduced to integrate high-level geographic features across multiple scales, enhancing feature representation prior to decoding. This work also highlights SAM's robust generalization capabilities with Digital Surface Model (DSM) data, a novel application. Extensive experiments on three benchmark multimodal remote sensing datasets, ISPRS Vaihingen, ISPRS Potsdam and MMHunan, demonstrate that the proposed MFNet significantly outperforms existing methods in multimodal semantic segmentation, setting a new standard in the field while offering a versatile foundation for future research and applications. The source code for this work is accessible at https://github.com/sstary/SSRS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。