用自监督学习提升无人机高分辨率多光谱图像的农业应用效果
Self-supervised training for high-resolution close-range multispectral remote sensing imagery
- 用MoCo-v3和掩码自编码器在多传感器多地区数据上预训练视觉模型
- 在德国和瑞士数据集上,预训练模型在少量标注数据下仍达领先精度
- 模型具备跨传感器、跨区域泛化能力,适合农业遥感小样本场景
尽管自监督学习(SSL)有望降低近距遥感的标注成本,但其在高分辨率多光谱无人机影像中的有效性仍受限于数据不足。本研究评估了基于厘米级多光谱无人机影像的自监督预训练在精准农业中的应用,数据涵盖多个传感器、年份和地区。采用Transformer编码器,在整合msuav500K与芬兰农田新采集的多年多光谱影像构建的统一数据集上,使用四个波段(绿、红、红边、近红外)进行预训练。在WeedMap数据集上以5%至100%训练数据进行作物杂草语义分割评估。两个下游任务分别为:任务A(德国,RedEdge-M),比较各预训练模型在部分与全微调下的表现;任务B(瑞士,Sequoia),验证任务A中最佳编码器的性能。基于MoCo-v3预训练的Swin Transformer在两任务中均表现最优,优于Doornbos等基于msuav500K预发布版训练的Swin模型。该模型还展现出良好的跨传感器与跨区域泛化能力。研究同时公开了芬兰多年度多光谱无人机数据集,以支持后续研究。
原文摘要 · Abstract (English)
Although self-supervised learning (SSL) offers a promising way to reduce annotation effort in close-range remote sensing, its effectiveness for high-resolution multispectral unmanned aerial vehicle (UAV) imagery remains underexplored due to limited data. This study evaluated SSL pretraining for precision agriculture using cm-scale multispectral drone imagery collected across multiple sensors, years, and regions. Transformer-based encoders were pretrained with Momentum Contrast v3 (MoCo-v3) and Masked Autoencoders on a harmonized dataset combining msuav500K with newly collected multi-year UAV imagery from agricultural fields in Finland. Pretraining used four spectral bands (Green, Red, Red-Edge, Near-Infrared) for cross-sensor compatibility. The models were evaluated on crop-weed semantic segmentation using the WeedMap dataset with 5--100% training data. The following two subsets served as downstream tasks: Task A (Germany, RedEdge-M), where all pretrained models were compared under partial and full fine-tuning, and Task B (Switzerland, Sequoia), where the best encoder from Task A was assessed. Our Swin Transformer pretrained with MoCo-v3 achieved the strongest performance on both tasks, surpassing the Swin Transformer model of Doornbos et al. pretrained on a pre-release of msuav500K. Our pretrained Swin Transformer further demonstrated cross-sensor and cross-region generalization. We additionally provide a public multi-year multispectral UAV dataset from Finland to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。