arXiv:2510.22665cs.CVcs.AI2025-10

首个专为遥感雷达图设计的图文基础模型,提升语义理解能力。

SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery

  • 构建百万级图文数据集,分两阶段迁移自然图像知识到雷达图域。
  • 在13个基准上超越现有模型,零样本分类准确率达82.4%。
  • 适合遥感、军事侦察等需要精准图像理解的领域使用。

合成孔径雷达(SAR)因其全天候成像能力至关重要。尽管自监督学习和掩码图像建模(MIM)已推动SAR基础模型发展,但现有方法多关注低层视觉特征,忽视多模态表征。此外,SAR多模态数据稀缺,制约跨模态模型发展。为此,我们构建了SARVLM-1M,一个包含超过一百万张图像-文本对的大规模视觉语言数据集,数据来自现有公开资源。为缓解SAR与自然图像间的显著差异,提出两阶段域迁移训练策略,利用光学遥感数据作为中间桥梁,实现从自然图像到SAR域的有效知识迁移。基于此策略,开发出首个面向SAR的视觉语言基础模型SARVLM,包含SARCLIP和SARCap。采用集成策略增强模型跨场景泛化能力。此外,SARDet与SARRot验证了该框架在目标检测任务中的有效性。在13个基准上的实验表明,SARVLM在图像-文本检索、目标识别、零样本分类、目标检测、语义定位及图像描述生成等方面均表现优异,持续优于当前最优视觉语言模型,显著推进了SAR图像的语义理解能力。代码与数据集将发布于https://github.com/KlayMa527/SARVLM.git。

原文摘要 · Abstract (English)

Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in self-supervised learning and masked image modeling (MIM) have enabled SAR foundation models, these approaches primarily focus on low-level visual features and often neglect multi-modal representation. Moreover, multimodal data for SAR is scarce, limiting the development of robust cross-modal models. To address this limitation, we construct SARVLM-1M, a large-scale vision-language dataset comprising over one million image-text pairs aggregated from existing datasets. Furthermore, to mitigate the substantial differences between SAR and natural imagery, we propose a two-stage domain transfer training strategy that leverages optical remote sensing data as an intermediate bridge, facilitating effective knowledge transfer from natural images to SAR domains. Based on this strategy, we develop SARVLM, the first vision-language foundation model tailored for SAR, consisting of SARCLIP and SARCap. In addition, an ensemble strategy is utilized to improve the cross-scene generalization capability of the model. Moreover, SARDet and SARRot further validate the capability of the proposed framework in object detection. Extensive experiments on 13 benchmarks across image-text retrieval, target recognition, zero-shot classification, object detection, semantic localization, and image captioning demonstrate the superior feature extraction and interpretation capabilities of SARVLM. It consistently outperforms state-of-the-art vision-language models and advances semantic understanding in SAR imagery. Code and datasets will be released on https://github.com/KlayMa527/SARVLM.git.

遥感图像视觉语言SAR基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。