首个面向多模态遥感影像的中文标注数据集,助力跨模态理解研究。
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning

- 构建包含SAR与多光谱图像的10米/20米分辨率人工标注数据集
- 在三种模态上测试发现可见光图像表现最优,雷达图像仍具挑战
- 提供特定模态提示可提升所有指标,适合遥感与多模态研究者
图像描述已成为计算机视觉重要任务,使模型能生成对视觉内容的自然语言描述。尽管已有自然图像和高分辨率光学遥感影像的数据集,但针对多模态卫星数据的描述数据集仍较稀缺,尤其在SAR影像和中等分辨率传感器方面。我们提出Sentinel2Cap,一个由人工标注的多模态图像描述数据集,包含10米和20米空间分辨率的Sentinel-1 SAR与Sentinel-2多光谱图像块,覆盖多样地表类型。所有描述均经人工创建并严格验证,确保语义准确性和语言质量。为评估该数据集,我们在三种图像模态(RGB、多光谱、SAR伪彩色)上使用Qwen3-VL-8B-Instruct模型进行零样本描述生成。结果显示,RGB图像表现最佳,而SAR图像对视觉语言模型仍具挑战性。提供模态特异性上下文提示能持续提升所有指标表现。这些发现揭示了多模态遥感图像描述的难点,并凸显了人工标注数据集在推动跨模态场景理解研究中的价值。所有材料已公开可用。
原文摘要 · Abstract (English)
Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While several datasets exist for natural images and high-resolution optical remote sensing imagery, the availability of captioning datasets for multimodal satellite data remains limited, particularly for SAR imagery and medium-resolution sensors. We introduce Sentinel2Cap, a human-annotated multimodal captioning dataset containing Sentinel-1 SAR and Sentinel-2 multi-spectral image patches at 10 m and 20 m spatial resolution with diverse land cover compositions. Captions are created manually and carefully validated to ensure both semantic accuracy and linguistic quality. To evaluate Sentinel2Cap, we perform a zero-shot captioning using the Qwen3-VL-8B-Instruct model across three image modalities: RGB, multi-spectral, and SAR pseudo-RGB representations. Results show that RGB images achieve the highest captioning performance, while SAR images remain more challenging for vision-language models. Providing modality-specific contextual prompts consistently improves performance across all metrics. These findings highlight both the challenges of multimodal remote sensing image captioning and the potential value of human-annotated datasets for advancing research in cross-modal scene understanding. All the material is publicly avaiable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。