arXiv:2605.30467cs.CV2026-05

为极地高分辨率遥感定制了新模型,提升基础设施等目标检测精度。

Clustering Guided Domain-Specific Pretrained Foundation Model for Very High-Resolution Arctic Remote Sensing

  • 用聚类筛选300万图像块,避免重复数据,保留场景多样性。
  • 在4个极地数据集上,平均F1分数提升5-8个百分点,最高达15%以上。
  • 适合做极地高分遥感的细粒度识别,如冰川、建筑等目标检测。

本研究通过结合多样性强的区域尺度图像筛选与掩码自编码器(MAE)自监督预训练,构建了一个面向北极地区的遥感基础模型(RSFM)。利用光谱和采集元数据描述符,在267TB的Vantor高分辨率影像中,通过可扩展的亲和传播聚类流程,筛选出约300万张图像块,旨在减少视觉重复或低信息区域的过度采样,同时保持研究区域内广泛的场景多样性。基于此优化数据集,使用领域适配的MAE重建目标对ViT-Large编码器进行预训练,生成适用于极地场景的特征权重。该预训练编码器被集成至现有位置感知检测与分割框架,并在四个手工标注的北极数据集上评估。相比ImageNet初始化的ViT-Large基线模型,北极域预训练在基础设施、冰漂物、临时营地和交通网络任务上的前景均值F1分数分别达到0.87、0.72、0.93和0.87,提升约5-8个百分点;且优于Prithvi-EO-2.0模型,最小提升也达15个百分点以上。结果表明,在固定架构与MAE目标的前提下,通过区域性数据分布优化,可产出更具迁移能力的极地专用编码器,适用于多种高分辨率遥感应用。

原文摘要 · Abstract (English)

This study introduces a novel Arctic-focused remote sensing foundation model (RSFM) by combining diversity-aware regional-scale image curation with masked autoencoder (MAE) self-supervised pretraining of a Vision Transformer (ViT) encoder for very-high-spatial-resolution (VHSR) satellite image analysis. Spectral and acquisition-metadata descriptors were used in a scalable affinity-propagation clustering workflow to select approximately 3 million chips from 267 TB of Vantor VHSR imagery This curation strategy was designed to reduce oversampling of visually repetitive or low-information areas while preserving broad scene diversity across the study domain. We pretrained a ViT-Large encoder on the curated corpus using a domain-adapted MAE reconstruction objective, producing Arctic-specific transformer weights for downstream feature mapping. The pretrained encoder was integrated into an existing location-aware detection and segmentation framework and evaluated across four hand-labeled Arctic datasets. Compared to ImageNet-initialized ViT-Large baseline, Arctic MAE pretraining produced consistent improvements in foreground mean F1 scores of 0.87, 0.72, 0.93, and 0.87, for infrastructure, IWP, RTS, and TCNs, with approximately 5-8 percentage increase. The proposed model also outperformed Prithvi-EO-2.0 in all downstream comparisons, with the smallest gain corresponding to at least a 15 percentage improvement mean F1, suggesting that domain-specific self-supervised pretraining on curated Arctic VHSR imagery provides more transferable representations for fine-scale Arctic mapping than a general-purpose Earth observation foundation model. These results demonstrate that optimizing the pretraining data distribution at regional scale, while keeping the architecture and MAE objective fixed, can produce a reusable Arctic-domain encoder for multiple VHSR remote sensing applications.

遥感极地监测自监督学习图像分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。