arXiv:2506.02868cs.CV2025-06被引 3

用视觉变压器+位置嵌入,精准识别北极冻土地貌与建筑。

Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings

  • 将位置信息融入视觉变压器,提升模型跨区域适应能力。
  • 在冻融滑塌检测上F1分数从0.84提升至0.92,优于传统模型。
  • 适合做高分辨率北极遥感分析的研究者和环保机构参考。

利用亚米级卫星影像在泛北极尺度上精确绘制冻土地貌、融冻扰动及人类基础设施,正变得日益重要。处理千兆级图像数据需高性能计算与鲁棒的特征检测模型。尽管卷积神经网络(CNN)在遥感中广泛应用,但类似大语言模型中成功应用的视觉变压器(ViT)通过注意力机制可更好捕捉长距离依赖与全局上下文。ViT支持自监督预训练,缓解北极标注数据稀缺问题,在基准数据集上表现优于CNN。北极场景还面临模型泛化挑战,同类语义特征呈现多样光谱特性。为此,本文在ViT中引入地理空间位置嵌入以增强区域适应性。研究评估了预训练ViT作为高分辨率北极遥感特征提取器的适用性,以及图像与位置嵌入融合的收益。基于已有北极特征检测数据集,对冰楔多边形(IWP)、 retrogressive thaw slumps(RTS)及人工基础设施三类任务进行验证。实验探索多种融合配置,结果表明加入位置嵌入的ViT在两项任务中超越先前基于CNN的模型,其中RTS检测的F1分数由0.84提升至0.92,证明具备空间感知能力的变压器模型在北极遥感应用中的潜力。

原文摘要 · Abstract (English)

Accurate mapping of permafrost landforms, thaw disturbances, and human-built infrastructure at pan-Arctic scale using sub-meter satellite imagery is increasingly critical. Handling petabyte-scale image data requires high-performance computing and robust feature detection models. While convolutional neural network (CNN)-based deep learning approaches are widely used for remote sensing (RS),similar to the success in transformer based large language models, Vision Transformers (ViTs) offer advantages in capturing long-range dependencies and global context via attention mechanisms. ViTs support pretraining via self-supervised learning-addressing the common limitation of labeled data in Arctic feature detection and outperform CNNs on benchmark datasets. Arctic also poses challenges for model generalization, especially when features with the same semantic class exhibit diverse spectral characteristics. To address these issues for Arctic feature detection, we integrate geospatial location embeddings into ViTs to improve adaptation across regions. This work investigates: (1) the suitability of pre-trained ViTs as feature extractors for high-resolution Arctic remote sensing tasks, and (2) the benefit of combining image and location embeddings. Using previously published datasets for Arctic feature detection, we evaluate our models on three tasks-detecting ice-wedge polygons (IWP), retrogressive thaw slumps (RTS), and human-built infrastructure. We empirically explore multiple configurations to fuse image embeddings and location embeddings. Results show that ViTs with location embeddings outperform prior CNN-based models on two of the three tasks including F1 score increase from 0.84 to 0.92 for RTS detection, demonstrating the potential of transformer-based models with spatial awareness for Arctic RS applications.

视觉变压器北极遥感位置嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。