构建首个大规模多传感器遥感图文数据集,助力地球观测AI理解复杂地表信息。
BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation
- 融合哨兵1号雷达与哨兵2号多光谱影像,生成960万条多样化文本标注。
- 在复杂土地利用分类任务中,现有视觉语言模型表现有限,微调后性能显著提升。
- 适合遥感、地球观测、多模态学习研究者使用,推动跨模态智能分析发展。
视觉-语言模型(VLMs)在计算机视觉领域表现优异,但在遥感(RS)数据上的表现仍受限于缺乏大规模、多传感器且文本标注丰富的遥感图像-文本数据集。现有数据集多为航空红绿蓝影像,文本描述短且弱关联,标注类型单一。为此,我们提出BigEarthNet.txt,一个大规模、多传感器的遥感图像-文本数据集,旨在推进地球观测中基于指令的图文学习。该数据集包含464,044对共注册的哨兵1号合成孔径雷达与哨兵2号多光谱图像,附带960万条文本标注,涵盖三类:i) 地理锚定的描述性文本,涵盖土地利用/覆被(LULC)类别、空间关系及环境上下文;ii) 适用于不同任务的视觉问答对;iii) 用于边界框预测的指代表达检测指令。通过统计对比,证明BigEarthNet.txt在文本丰富度和标注多样性上优于现有遥感图文数据集。我们进一步构建了人工验证的基准划分,用于评估VLM在遥感与计算机视觉任务中的表现。结果表明,现有模型在涉及复杂LULC类别的任务中存在局限,而使用BigEarthNet.txt进行微调后,所有任务性能均实现持续提升。
原文摘要 · Abstract (English)
Vision-langugage models (VLMs) have shown strong performance in computer vision (CV), yet their performance on remote sensing (RS) data remains limited due to the lack of large-scale, multi-sensor RS image-text datasets with diverse textual annotations. Existing datasets predominantly include aerial Red-Green-Blue imagery, with short or weakly grounded captions, and provide limited diversity in annotation types. To address this limitation, we introduce BigEarthNet$.$txt, a large-scale, multi-sensor image-text dataset designed to advance instruction-driven image-text learning in Earth observation across multiple tasks. BigEarthNet$.$txt contains 464044 co-registered Sentinel-1 synthetic aperture radar and Sentinel-2 multispectral images with 9.6M text annotations, including: i) geographically anchored captions describing land-use/land-cover (LULC) classes, their spatial relations, and environmental context; ii) visual question answering pairs relevant for different tasks; and iii) referring expression detection instructions for bounding box prediction. Through a comparative statistical analysis, we demonstrate that BigEarthNet$.$txt surpasses existing RS image-text datasets in textual richness and annotation type variety. We further establish a manually-verified benchmark split to evaluate VLMs in RS and CV. The results show the limitations of these models on tasks that involve complex LULC classes, whereas fine-tuning using BigEarthNet$.$txt results in consistent performance gains across all considered tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。