解决超高分辨率卫星图像中类别分布不均问题,提升小类分割精度。
SRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation
- 采用多尺度裁剪与语义重排重采样增强数据
- 首次无区域文本描述融合图文特征,提升模型泛化能力
- 在三个数据集上显著提升分割精度,适合遥感图像细分任务
长尾分布严重制约了超分辨率(UHR)卫星影像语义分割的发展。现有方法多聚焦于多分支网络的多尺度特征提取与融合,却忽视了长尾问题。本文提出SRMF框架,通过多尺度裁剪结合语义重排重采样策略进行数据增强,并创新性地提出无需区域文本描述的图文特征融合方法,注入通用表征知识以增强模型鲁棒性。在URUR、GID和FBP数据集上的实验表明,该方法分别提升mIoU 3.33%、0.66%和0.98%,达到当前最优性能。
原文摘要 · Abstract (English)
The long-tail problem presents a significant challenge to the advancement of semantic segmentation in ultra-high-resolution (UHR) satellite imagery. While previous efforts in UHR semantic segmentation have largely focused on multi-branch network architectures that emphasize multi-scale feature extraction and fusion, they have often overlooked the importance of addressing the long-tail issue. In contrast to prior UHR methods that focused on independent feature extraction, we emphasize data augmentation and multimodal feature fusion to alleviate the long-tail problem. In this paper, we introduce SRMF, a novel framework for semantic segmentation in UHR satellite imagery. Our approach addresses the long-tail class distribution by incorporating a multi-scale cropping technique alongside a data augmentation strategy based on semantic reordering and resampling. To further enhance model performance, we propose a multimodal fusion-based general representation knowledge injection method, which, for the first time, fuses text and visual features without the need for individual region text descriptions, extracting more robust features. Extensive experiments on the URUR, GID, and FBP datasets demonstrate that our method improves mIoU by 3.33\%, 0.66\%, and 0.98\%, respectively, achieving state-of-the-art performance. Code is available at: https://github.com/BinSpa/SRMF.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。