构建了首个细粒度合成孔径雷达图像描述数据集,助力智能遥感分析。
FSAR-Cap: A Fine-Grained Two-Stage Annotated Dataset for SAR Image Captioning
- 采用两阶段标注策略,结合模板生成与人工校验提升质量
- 包含14,480张图像和72,400对图文数据,覆盖更广类别
- 适用于遥感、军事侦察等领域的图像理解研究
合成孔径雷达(SAR)图像描述能实现场景级语义理解,在军事情报和城市规划等领域具有重要作用,但其发展受限于高质量数据集的匮乏。为此,我们提出FSAR-Cap,一个大规模SAR图像描述数据集,包含14,480张图像和72,400对图像-文本样本。该数据集基于FAIR-CSAR目标检测数据集构建,采用两阶段标注策略,结合分层模板表示、人工验证与补充及提示标准化。相比现有资源,FSAR-Cap提供更丰富的细粒度标注、更广泛的类别覆盖和更高的标注质量。通过多种编码器-解码器架构的基准测试,验证了其有效性,为未来SAR图像描述与智能图像解析研究奠定了基础。
原文摘要 · Abstract (English)
Synthetic Aperture Radar (SAR) image captioning enables scene-level semantic understanding and plays a crucial role in applications such as military intelligence and urban planning, but its development is limited by the scarcity of high-quality datasets. To address this, we present FSAR-Cap, a large-scale SAR captioning dataset with 14,480 images and 72,400 image-text pairs. FSAR-Cap is built on the FAIR-CSAR detection dataset and constructed through a two-stage annotation strategy that combines hierarchical template-based representation, manual verification and supplementation, prompt standardization. Compared with existing resources, FSAR-Cap provides richer fine-grained annotations, broader category coverage, and higher annotation quality. Benchmarking with multiple encoder-decoder architectures verifies its effectiveness, establishing a foundation for future research in SAR captioning and intelligent image interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。