arXiv:2507.18743cs.CV2025-07被引 7

构建13万对SAR遥感图文数据集,提升雷达图像语义理解能力。

SAR-TEXT: A Large-Scale SAR Image-Text Dataset Built with SAR-Narrator and A Progressive Learning Strategy for Downstream Tasks

  • 用多阶段框架SAR-Narrator自动生成高质量图文对
  • 在检索、描述生成和问答任务上性能提升10%以上
  • 开源数据集与模型,适合遥感与多模态研究者使用

近年来,视觉语言模型在遥感领域取得显著进展。合成孔径雷达(SAR)影像具备全天候成像能力,在遥感中至关重要,但缺乏大规模高质量的SAR图像-文本数据集,制约了其语义理解。本文构建了SAR-TEXT,一个包含超过13万对SAR图像-文本的数据集。为构建该数据集,设计了SAR-Narrator框架,通过多阶段策略生成图像文本描述。为验证数据集有效性,我们在三类典型视觉语言任务上开展实验:图像-文本检索、图像描述生成和视觉问答(VQA)。具体构建了三种代表性模型:SAR-RS-CLIP、SAR-RS-CoCa和SAR-GPT。SAR-RS-CLIP在检索任务中表现突出,使OSdataset_512和HRSID测试集平均召回率分别提升12.97%和10.0%。在描述生成任务中,SAR-RS-CoCa在BLEU-4、SPICE和CIDEr指标上显著优于原始CoCa模型。在VQA任务中,SAR-GPT在多个SAR-VQA数据集上超越基线与单阶段模型,展现出更强的语义理解与推理能力,定性结果进一步验证。值得注意的是,SAR-Narrator作为灵活的描述生成工具,可被社区用于构建更大规模的SAR图文数据集。所有代码、预训练模型及SAR-TEXT数据集均公开于https://github.com/YiguoHe/SAR-TEXT。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) have achieved remarkable breakthroughs in the field of remote sensing in recent years. Synthetic Aperture Radar (SAR) imagery, with its all-weather capability, is essential in remote sensing, yet the lack of large-scale, high-quality SAR image-text datasets hinders its semantic understanding. In this paper, we construct SAR-TEXT, a large-scale and high-quality dataset consisting of over 130,000 SAR image-text pairs. To construct the SAR-TEXT dataset, we design the SAR-Narrator framework, which generates textual descriptions for SAR images through a multi-stage strategy. To verify the effectiveness of the SAR-TEXT dataset, we conduct experiments on three typical vision-language tasks: image-text retrieval, image captioning, and visual question answering (VQA). Specifically, we construct three representative models on SAR-TEXT: SAR-RS-CLIP, SAR-RS-CoCa, and SAR-GPT. SAR-RS-CLIP achieves notable improvements in retrieval performance, boosting average recall by 12.97% and 10.0% on the OSdataset_512 and HRSID test sets, respectively. In the captioning task, SAR-RS-CoCa achieves significant improvements over the original CoCa models in terms of BLEU-4, SPICE, and CIDEr scores. In the VQA task, SAR-GPT outperforms baseline and single-stage models on multiple SAR-VQA datasets, demonstrating stronger semantic understanding and reasoning ability, as further confirmed by qualitative results. It is worth noting that, as a flexible captioning tool, SAR-Narrator can be readily adopted by the community to construct larger-scale SAR image-text datasets. All code, pretrained models, and the SAR-Text dataset are publicly available at: https://github.com/YiguoHe/SAR-TEXT.

遥感图文生成SAR数据集多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。