arXiv:2412.14428cs.CVcs.LG2024-12ICCV被引 15

用野生动物观测数据训练卫星图像模型,提升遥感识别与跨模态检索能力。

WildSAT: Learning Satellite Image Representations from Wildlife Observations

  • 通过对比学习融合卫星图、物种分布和文本描述,构建多模态表示
  • 在多个遥感任务上超越ImageNet预训练模型,零样本检索效果优异
  • 适合遥感、生物多样性监测领域研究者使用,可扩展至跨模态应用

物种分布蕴含丰富的生态与环境信息,但其在遥感表征学习中的潜力尚未被充分挖掘。我们提出WildSAT,将数百万地理标注的野生动物观测数据与卫星图像配对,利用对比学习联合优化卫星图像、物种出现地图及文本栖息地描述,实现模型训练或微调。该方法显著提升多种卫星图像识别任务性能,优于ImageNet预训练模型与专用遥感基线。通过视觉与文本对齐,WildSAT支持零样本检索,用户可基于文本描述搜索地理区域。其表现超越近期跨模态学习方法,包括卫星图像与地面影像或野生动物照片对齐的方案,凸显本方法优势。最后,我们分析关键设计选择,强调其在遥感与生物多样性监测中的广泛适用性。

原文摘要 · Abstract (English)

Species distributions encode valuable ecological and environmental information, yet their potential for guiding representation learning in remote sensing remains underexplored. We introduce WildSAT, which pairs satellite images with millions of geo-tagged wildlife observations readily-available on citizen science platforms. WildSAT employs a contrastive learning approach that jointly leverages satellite images, species occurrence maps, and textual habitat descriptions to train or fine-tune models. This approach significantly improves performance on diverse satellite image recognition tasks, outperforming both ImageNet-pretrained models and satellite-specific baselines. Additionally, by aligning visual and textual information, WildSAT enables zero-shot retrieval, allowing users to search geographic locations based on textual descriptions. WildSAT surpasses recent cross-modal learning methods, including approaches that align satellite images with ground imagery or wildlife photos, demonstrating the advantages of our approach. Finally, we analyze the impact of key design choices and highlight the broad applicability of WildSAT to remote sensing and biodiversity monitoring.

遥感跨模态生态学对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。