首个专为SAR图像设计的视觉语言模型,提升全天候遥感理解能力。
FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery
- 引入地理基线与时空锚点,融合多源遥感时间特征增强视觉表征
- 在多个基准测试中超越主流模型超10%,达到当前最佳性能
- 适合遥感、气象、灾害监测等需全天候成像理解的研究者使用
全天气、全天时合成孔径雷达(SAR)的智能解译研究对推进遥感应用至关重要。近年来,尽管视觉语言模型(VLMs)在可见光图像上展现出强大的开放世界理解能力,但因其成像机制复杂、对散射特征敏感且高质量图文语料稀缺,在直接应用于SAR领域时性能严重受限。为此,我们构建了首个SAR图像-文本-AlphaEarth特征三元组数据集,并开发了专用于SAR的FUSAR-GPT模型。该模型创新性地引入地理基线模型作为‘世界知识’先验,并通过‘时空锚点’将多源遥感时间特征嵌入视觉主干网络,实现对SAR图像中目标稀疏表征的动态补偿。此外,设计了两阶段监督微调策略,解耦大模型的知识注入与任务执行过程。时空特征嵌入与两阶段解耦范式使FUSAR-GPT在多个典型遥感视觉-语言基准测试中均达到领先水平,显著优于主流基线模型超过10%。
原文摘要 · Abstract (English)
Research on the intelligent interpretation of all-weather, all-time Synthetic Aperture Radar (SAR) is crucial for advancing remote sensing applications. In recent years, although Visual Language Models (VLMs) have demonstrated strong open-world understanding capabilities on RGB images, their performance is severely limited when directly applied to the SAR field due to the complexity of the imaging mechanism, sensitivity to scattering features, and the scarcity of high-quality text corpora. To systematically address this issue, we constructed the inaugural SAR Image-Text-AlphaEarth feature triplet dataset and developed FUSAR-GPT, a VLM specifically for SAR. FUSAR-GPT innovatively introduces a geospatial baseline model as a 'world knowledge' prior and embeds multi-source remote-sensing temporal features into the model's visual backbone via 'spatiotemporal anchors', enabling dynamic compensation for the sparse representation of targets in SAR images. Furthermore, we designed a two-stage SFT strategy to decouple the knowledge injection and task execution of large models. The spatiotemporal feature embedding and the two-stage decoupling paradigm enable FUSAR-GPT to achieve state-of-the-art performance across several typical remote sensing visual-language benchmark tests, significantly outperforming mainstream baseline models by over 10%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。