arXiv:2412.08536cs.CV2024-12中稿 · WACV'25被引 19

用地面图像提升卫星图分类,让零样本土地利用识别更准

SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting

  • 用配对的地面与卫星图像微调CLIP,建立跨模态关联
  • 在EuroSAT和BigEarthNet上准确率提升显著,通用提示效果更好
  • 适合遥感、地理信息领域研究者,尤其关注零样本分类应用

预训练视觉语言模型(如CLIP)在自由文本提示下展现出出色的零样本分类能力,甚至在特定领域有一定泛化性。然而,由于其训练数据以地面图像为主,卫星影像代表性不足,导致在卫星图像上的表现受限。现有卫星图像提示方法多依赖“一张卫星图像的...”等通用短语,限制了零样本土地利用与土地覆盖(LULC)分类的效果。为此,我们提出SenCLIP,通过使用欧洲范围内大规模配对的Sentinel-2影像与地理标签地面照片数据集,将CLIP的表征迁移至Sentinel-2影像。我们在EuroSAT和BigEarthNet数据集上评估了SenCLIP在零样本LULC映射任务中的表现,对比了航空与地面提示风格。结果表明,通过将地面表征与卫星影像对齐,本方法在两种提示风格下均显著提升了分类准确率,为零样本LULC映射中使用自由文本描述开辟了新可能。

原文摘要 · Abstract (English)

Pre-trained vision-language models (VLMs), such as CLIP, demonstrate impressive zero-shot classification capabilities with free-form prompts and even show some generalization in specialized domains. However, their performance on satellite imagery is limited due to the underrepresentation of such data in their training sets, which predominantly consist of ground-level images. Existing prompting techniques for satellite imagery are often restricted to generic phrases like a satellite image of ..., limiting their effectiveness for zero-shot land-use and land-cover (LULC) mapping. To address these challenges, we introduce SenCLIP, which transfers CLIPs representation to Sentinel-2 imagery by leveraging a large dataset of Sentinel-2 images paired with geotagged ground-level photos from across Europe. We evaluate SenCLIP alongside other SOTA remote sensing VLMs on zero-shot LULC mapping tasks using the EuroSAT and BigEarthNet datasets with both aerial and ground-level prompting styles. Our approach, which aligns ground-level representations with satellite imagery, demonstrates significant improvements in classification accuracy across both prompt styles, opening new possibilities for applying free-form textual descriptions in zero-shot LULC mapping.

遥感零样本视觉语言模型土地利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。