arXiv:2502.08168cs.CL2025-02被引 14

首个面向雷达图像的多任务图文对话数据集,助力专业遥感理解

SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation

  • 构建200万条高质量雷达图像与文本配对数据,覆盖多样场景
  • 在16个主流模型上验证有效,推动视觉语言模型在遥感领域应用
  • 适合遥感、AI多模态研究者,为垂直领域数据构建提供范式

合成孔径雷达(SAR)是一种强大的全天候地球观测工具,广泛应用于军事侦察、海上监视和基础设施监测。尽管视觉语言模型(VLMs)在自然语言处理和图像理解方面取得显著进展,但在专业领域因缺乏领域知识而应用受限。本文首次提出面向SAR图像的大型多模态对话数据集SARChat-2M,包含约200万条高质量图像-文本对,涵盖多样化场景并配有详细目标标注。该数据集不仅支持视觉理解、目标检测等关键任务,还创新性地构建了首个SAR领域的视觉语言数据集与评测基准,可评估和推动VLM在SAR图像解读中的能力,为构建各类遥感垂直领域多模态数据集提供范式框架。通过在16个主流VLM上的实验,验证了数据集的有效性。项目将开源于https://github.com/JimmyMa99/SARChat。

原文摘要 · Abstract (English)

As a powerful all-weather Earth observation tool, synthetic aperture radar (SAR) remote sensing enables critical military reconnaissance, maritime surveillance, and infrastructure monitoring. Although Vision language models (VLMs) have made remarkable progress in natural language processing and image understanding, their applications remain limited in professional domains due to insufficient domain expertise. This paper innovatively proposes the first large-scale multimodal dialogue dataset for SAR images, named SARChat-2M, which contains approximately 2 million high-quality image-text pairs, encompasses diverse scenarios with detailed target annotations. This dataset not only supports several key tasks such as visual understanding and object detection tasks, but also has unique innovative aspects: this study develop a visual-language dataset and benchmark for the SAR domain, enabling and evaluating VLMs' capabilities in SAR image interpretation, which provides a paradigmatic framework for constructing multimodal datasets across various remote sensing vertical domains. Through experiments on 16 mainstream VLMs, the effectiveness of the dataset has been fully verified. The project will be released at https://github.com/JimmyMa99/SARChat.

遥感视觉语言多模态雷达图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。