arXiv:2412.02044eess.IVcs.CV2024-12被引 60

提出不对称对齐网络,提升遥感图像融合分类效果

ASANet: Asymmetric Semantic Aligning Network for RGB and SAR image land cover classification

  • 设计异构特征对齐机制,分别加权处理可见光与雷达图像
  • 在三个数据集上实现最高准确率,新数据集上提升达17.69%
  • 适合遥感图像融合、多模态分类研究者使用

合成孔径雷达(SAR)图像与可见光(RGB)图像结合,在多模态地表覆盖分类中具有重要价值。现有方法通常假设两模态特征一致,忽视各自特性。本文提出异步语义对齐网络(ASANet),在特征层面引入非对称性,以充分挖掘互补信息。核心为语义聚焦模块(SFM),显式计算各模态差异权重;同时引入级联融合模块(CFM),从通道与空间维度精细选择融合特征。二者协同使模型有效学习跨模态关联,抑制特征差异带来的噪声。实验表明,ASANet在三个多模态数据集上表现优异,并在新构建的RGB-SAR数据集上超越主流方法1.21%~17.69%。当输入图为256×256像素时,运行速度达48.7帧每秒。

原文摘要 · Abstract (English)

Synthetic Aperture Radar (SAR) images have proven to be a valuable cue for multimodal Land Cover Classification (LCC) when combined with RGB images. Most existing studies on cross-modal fusion assume that consistent feature information is necessary between the two modalities, and as a result, they construct networks without adequately addressing the unique characteristics of each modality. In this paper, we propose a novel architecture, named the Asymmetric Semantic Aligning Network (ASANet), which introduces asymmetry at the feature level to address the issue that multi-modal architectures frequently fail to fully utilize complementary features. The core of this network is the Semantic Focusing Module (SFM), which explicitly calculates differential weights for each modality to account for the modality-specific features. Furthermore, ASANet incorporates a Cascade Fusion Module (CFM), which delves deeper into channel and spatial representations to efficiently select features from the two modalities for fusion. Through the collaborative effort of these two modules, the proposed ASANet effectively learns feature correlations between the two modalities and eliminates noise caused by feature differences. Comprehensive experiments demonstrate that ASANet achieves excellent performance on three multimodal datasets. Additionally, we have established a new RGB-SAR multimodal dataset, on which our ASANet outperforms other mainstream methods with improvements ranging from 1.21% to 17.69%. The ASANet runs at 48.7 frames per second (FPS) when the input image is 256x256 pixels. The source code are available at https://github.com/whu-pzhang/ASANet

多模态遥感图像特征融合深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。