面向安第斯地区多光谱遥感的自监督视觉大模型,提升考古特征识别效率
DeepAndes: A Self-Supervised Vision Foundation Model for Multi-Spectral Remote Sensing Imagery of the Andes
- 基于三百万张多光谱影像,用定制DINOv2自监督训练
- 少样本场景下分类、检索、分割任务性能超越基线模型
- 首个专为安第斯考古设计的视觉基础模型,适合遥感考古研究者
通过大规模遥感数据映射遗址,考古学家可洞察长期人口趋势、区域社会网络及过去对气候变化的适应。遥感调查结合深度学习与计算机视觉技术,能显著扩展研究范围。然而,传统监督学习在细粒度考古特征标注上面临挑战。尽管近期视觉基础模型在少标注条件下表现优异,但多数现成方案仅针对RGB图像,不适用于8波段多光谱数据。本文提出DeepAndes,一种基于Transformer的视觉基础模型,专门针对安第斯地区考古应用,使用三百万张8波段多光谱卫星图像进行训练。该模型采用定制化的DINOv2自监督学习算法,是首个专为安第斯地区设计的基础模型。我们在不平衡图像分类、图像实例检索和像素级语义分割任务上评估其性能,结果显示,在少样本学习场景中,DeepAndes在F1分数、平均精度和Dice系数上均显著优于从头训练或在小数据集上预训练的模型,证明了大规模自监督预训练在考古遥感中的有效性。代码将公开于https://github.com/geopacha/DeepAndes。
原文摘要 · Abstract (English)
By mapping sites at large scales using remotely sensed data, archaeologists can generate unique insights into long-term demographic trends, inter-regional social networks, and past adaptations to climate change. Remote sensing surveys complement field-based approaches, and their reach can be especially great when combined with deep learning and computer vision techniques. However, conventional supervised deep learning methods face challenges in annotating fine-grained archaeological features at scale. While recent vision foundation models have shown remarkable success in learning large-scale remote sensing data with minimal annotations, most off-the-shelf solutions are designed for RGB images rather than multi-spectral satellite imagery, such as the 8-band data used in our study. In this paper, we introduce DeepAndes, a transformer-based vision foundation model trained on three million multi-spectral satellite images, specifically tailored for Andean archaeology. DeepAndes incorporates a customized DINOv2 self-supervised learning algorithm optimized for 8-band multi-spectral imagery, marking the first foundation model designed explicitly for the Andes region. We evaluate its image understanding performance through imbalanced image classification, image instance retrieval, and pixel-level semantic segmentation tasks. Our experiments show that DeepAndes achieves superior F1 scores, mean average precision, and Dice scores in few-shot learning scenarios, significantly outperforming models trained from scratch or pre-trained on smaller datasets. This underscores the effectiveness of large-scale self-supervised pre-training in archaeological remote sensing. Codes will be available on https://github.com/geopacha/DeepAndes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。