构建首个覆盖36年澳大利亚遥感影像的多卫星视觉语言数据集
Landsat30-AU: A Vision-Language Dataset for Australian Landsat Imagery
- 基于四颗陆地卫星30米分辨率影像,构建大规模图文对与视觉问答数据
- 现有模型在遥感理解上表现差,微调后性能显著提升
- 适合遥感、地球观测和多模态学习研究者使用
视觉语言模型(VLMs)可实现自然语言与卫星图像的交互,推动地球观测民主化,加速专家工作流,让非专业人士也能使用数据,并支持全球尺度自动化。然而,现有数据集主要聚焦短期、高分辨率卫星影像,忽视了低分辨率、多卫星、长期存档的兰斯特定(Landsat)数据,而这正是实现低成本、抗偏差全球监测的关键。为此,我们提出兰斯特定30-AU(Landsat30-AU),一个基于四颗陆地卫星(5、7、8、9)在澳大利亚超过36年采集的30米分辨率影像构建的大规模视觉语言数据集。该数据集包含两个部分:Landsat30-AU-Cap,含196,262张图像-标题对;Landsat30-AU-VQA,含17,725个经人工验证的跨八个遥感领域的视觉问答样本。数据通过自举式流程生成,结合通用VLM迭代优化与人工校验以保障质量。我们在该基准上评估八种VLM,发现现成模型难以理解卫星图像:开源遥感VLM EarthDial在图像描述任务中仅得0.07 SPIDEr,VQA准确率为0.48。令人鼓舞的是,对Qwen2.5-VL-7B进行轻量级微调后,其描述得分从0.11升至0.31 SPIDEr,VQA准确率从0.74升至0.87。代码与数据已公开于https://github.com/papersubmit1/landsat30-au。
原文摘要 · Abstract (English)
Vision language models (VLMs) that enable natural language interaction with satellite imagery can democratize Earth observation by accelerating expert workflows, making data accessible to non-specialists, and enabling planet-scale automation. However, existing datasets focus mainly on short-term, high-resolution imagery from a limited number of satellites, overlooking low-resolution, multi-satellite, long-term archives, such as Landsat, that are essential for affordable and bias-robust global monitoring. We address this gap with Landsat30-AU, a large-scale vision-language dataset built from 30-meter resolution imagery collected by four Landsat satellites (5, 7, 8, and 9) over Australia, spanning more than 36 years. The dataset includes two components: Landsat30-AU-Cap, containing $196,262$ image-caption pairs, and Landsat30-AU-VQA, comprising 17,725 human-verified visual question answering (VQA) samples across eight remote sensing domains. Both datasets are curated through a bootstrapped pipeline that leverages generic VLMs with iterative refinement and human verification to ensure quality. Our evaluation of eight VLMs on our benchmark reveals that off-the-shelf models struggle to understand satellite imagery. The open-source remote-sensing VLM EarthDial achieves only 0.07 SPIDEr in captioning and a VQA accuracy of 0.48, highlighting the limitations of current approaches. Encouragingly, lightweight fine-tuning of Qwen2.5-VL-7B on Landsat30-AU improves captioning performance from 0.11 to 0.31 SPIDEr and boosts VQA accuracy from 0.74 to 0.87. Code and data are available at https://github.com/papersubmit1/landsat30-au.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。