arXiv:2509.23927cs.CV2025-09被引 7

首个面向SAR遥感的多模态大模型,融合地理知识提升理解能力。

FUSAR-KLIP: Towards Multimodal Foundation Models for Remote Sensing

  • 构建首个带地理投影的百万级SAR数据集,覆盖135城12万图
  • 用分层认知链生成结构化文本,编码上百万维地理语义信息
  • 自洽迭代优化机制让模型学习符合人类认知与物理规律

通用视觉语言模型在图像理解上取得突破,但通用视觉表征与遥感图像解读存在根本认知差异:遥感图像蕴含地形、地貌与空间结构信息,需深度地学理解。合成孔径雷达(SAR)影像更因相干成像机制,与普通图像呈现显著模态异质性。为弥合此差异,本文提出FUSAR-KLIP——首个面向SAR的知识引导多模态基础模型,并提供可复用数据与评估基准。具体包括:(1) 构建FUSAR-GEOVL-1M,首个含完整地理投影属性的大规模SAR数据集,覆盖多个卫星平台,含12万张图像和135个城市;(2) 通过分层认知思维链生成结构化文本,精准编码超过100万条多维度语义信息,涵盖地貌环境、区域属性及空间关系;(3) 设计自洽迭代优化机制,在对比、匹配与重建构成的自监督闭环中,引导模型学习符合人类认知与物理法则的跨模态表示;(4) 在两大类共11个典型下游任务中建立统一评估基准,并与15个主流基础模型对比。

原文摘要 · Abstract (English)

Cross-modal artificial intelligence, represented by visual language models, has achieved significant success in general image understanding. However, a fundamental cognitive inconsistency exists between general visual representation and remote sensing image interpretation: remote sensing images couple topography, terrain, and spatial structure, thereby inherently requiring models to possess deep geoscientific understanding. This cognitive difference is further amplified in synthetic aperture radar (SAR) imagery: while SAR possesses irreplaceable all-weather, all-day observation capabilities, it is constrained by coherent imaging mechanisms, exhibiting significant modal heterogeneity with general images. To address this inconsistency, we propose FUSAR-KLIP, the first knowledge-guided general multimodal foundational model for SAR, along with reusable data and evaluation baselines. Specifically: (1) FUSAR-GEOVL-1M (the first large-scale SAR dataset with complete geographic projection attributes) was constructed, covering multiple satellite platforms, 120,000 images, and 135 cities; (2) Aligned structured text was generated through hierarchical cognitive thought chains, accurately encoding more than 1 million multidimensional semantic information from geomorphological environment and regional attributes to spatial relationships; (3) A self-consistent iterative optimization mechanism was designed to guide cross-modal learning with this knowledge information consistent with human cognition and physical laws in a self-supervised closed loop consisting of contrast, matching, and reconstruction; (4) A unified evaluation benchmark was established in 11 typical downstream tasks in the two major categories of vision and language, and compared with 15 mainstream foundation models.

遥感多模态SAR大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。