arXiv:2511.10774cs.CV2025-11被引 1

提出频域感知的遥感多模态分类模型,提升跨场景泛化能力。

Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification

  • 设计频域解耦模块,分离高低频特征以应对多模态差异。
  • 在多个遥感数据集上超越当前最佳方法,准确率提升2.1%以上。
  • 适合需要跨传感器、跨场景泛化的遥感图像分析任务。

遥感技术的快速发展催生了多模态泛化新任务,要求模型克服数据异构性并具备强跨场景泛化能力。现有视觉-语言模型通常使用通用文本描述地表材料,缺乏针对不同遥感模态的专有语言先验。本文将遥感多模态泛化(RSMG)形式化为学习范式,提出频域感知的视觉-语言多模态泛化网络(FVMGN)。设计基于扩散的训练-测试时增强(DTAug)策略,重建多模态地表覆盖分布,丰富输入信息。提出多模态小波解耦(MWDis)模块,在频域重采样高低频成分,学习跨域不变特征。针对遥感模态特性,设计共享与专属类别文本作为变压器文本编码器输入,提取多样化文本特征。构建空间-频域感知图像编码器(SFIE),实现局部-全局特征重建与表征。最后,引入多尺度空间-频域特征对齐(MSFFA)模块,在空间与频域构建统一语义空间,实现文本与视觉特征的精细化对齐。大量实验表明,相比主流方法,FVMGN在多模态泛化能力上表现卓越,多个数据集上准确率提升2.1%以上。

原文摘要 · Abstract (English)

The booming remote sensing (RS) technology is giving rise to a novel multimodality generalization task, which requires the model to overcome data heterogeneity while possessing powerful cross-scene generalization ability. Moreover, most vision-language models (VLMs) usually describe surface materials in RS images using universal texts, lacking proprietary linguistic prior knowledge specific to different RS vision modalities. In this work, we formalize RS multimodality generalization (RSMG) as a learning paradigm, and propose a frequency-aware vision-language multimodality generalization network (FVMGN) for RS image classification. Specifically, a diffusion-based training-test-time augmentation (DTAug) strategy is designed to reconstruct multimodal land-cover distributions, enriching input information for FVMGN. Following that, to overcome multimodal heterogeneity, a multimodal wavelet disentanglement (MWDis) module is developed to learn cross-domain invariant features by resampling low and high frequency components in the frequency domain. Considering the characteristics of RS vision modalities, shared and proprietary class texts is designed as linguistic inputs for the transformer-based text encoder to extract diverse text features. For multimodal vision inputs, a spatial-frequency-aware image encoder (SFIE) is constructed to realize local-global feature reconstruction and representation. Finally, a multiscale spatial-frequency feature alignment (MSFFA) module is suggested to construct a unified semantic space, ensuring refined multiscale alignment of different text and vision features in spatial and frequency domains. Extensive experiments show that FVMGN has the excellent multimodality generalization ability compared with state-of-the-art (SOTA) methods.

遥感图像多模态频域分析视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。