arXiv:2511.03004cs.CV2025-11被引 2

仅用1000个标注样本,实现1米分辨率全省土地覆盖分类。

Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning

  • 利用自监督预训练减少对标注数据依赖
  • 整体准确率87.14%,宏F1达75.58%
  • 适合大规模高分辨率土地覆盖制图场景

深度学习语义分割在1米分辨率土地覆盖分类中表现优异,但大规模代表性训练数据的获取成为广泛应用的主要障碍。本研究提出一种标签高效的新型方法,仅使用1,000个标注参考图像块,结合自监督深度学习完成全州1米分辨率土地覆盖分类。采用“自举你的潜在表示”(BYOL)策略,利用377,921张未标注的彩色红外航空影像(256×256像素,1米分辨率)对ResNet-101卷积编码器进行预训练。将学习到的编码器权重迁移至多种语义分割架构(FCN、U-Net、Attention U-Net、DeepLabV3+、UPerNet、PAN),并在极小训练集(250、500、750块)上通过交叉验证微调。最终集成最优U-Net模型,在覆盖美国密西西比州超1230亿像素的8类土地覆盖映射中达到87.14%的整体准确率和75.58%的宏F1分数。定性与定量分析表明,水体与林地分类准确,但耕地、草本及裸地之间的区分仍具挑战。结果表明,自监督学习可有效降低对大量人工标注数据的需求,直接缓解高分辨率土地覆盖制图规模化应用的核心瓶颈。

原文摘要 · Abstract (English)

Deep learning semantic segmentation methods have shown promising performance for very high 1-m resolution land cover classification, but the challenge of collecting large volumes of representative training data creates a significant barrier to widespread adoption of such models for meter-scale land cover mapping over large areas. In this study, we present a novel label-efficient approach for statewide 1-m land cover classification using only 1,000 annotated reference image patches with self-supervised deep learning. We use the "Bootstrap Your Own Latent" pre-training strategy with a large amount of unlabeled color-infrared aerial images (377,921 patches of 256x256 pixels at 1-m resolution) to pre-train a ResNet-101 convolutional encoder. The learned encoder weights were subsequently transferred into multiple deep semantic segmentation architectures (FCN, U-Net, Attention U-Net, DeepLabV3+, UPerNet, PAN), which were then fine-tuned using very small training dataset sizes with cross-validation (250, 500, 750 patches). Among the fine-tuned models, we obtained 87.14% overall accuracy and 75.58% macro F1 score using an ensemble of the best-performing U-Net models for comprehensive 1-m, 8-class land cover mapping, covering more than 123 billion pixels over the state of Mississippi, USA. Detailed qualitative and quantitative analysis revealed accurate mapping of open water and forested areas, while highlighting challenges in accurate delineation between cropland, herbaceous, and barren land cover types. These results show that self-supervised learning is an effective strategy for reducing the need for large volumes of manually annotated data, directly addressing a major limitation to high spatial resolution land cover mapping at scale.

土地覆盖自监督学习高分辨率语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。