用轻量级结构让遥感大模型更准识别休耕土地,助力粮食水资源协同管理
Adapting Prithvi-EO for Fallow Detection for Food-Water Nexus: ViT-Adapter Necks and Parameter-Efficient Backbone tuning of Geospatial Foundation Model

- 设计轻量级ViT-Adapter颈部,融合多尺度空间先验信息
- 采用低秩适配与选择性微调,使模型在休耕检测上达mAP@50 0.9479
- 适合关注农业遥感、资源管理的科研与政策应用者
理解休耕土地的空间分布对优化粮食-水资源协同关系至关重要,因休耕在轮作和节水中的作用。然而,休耕在美农部耕地数据层(CDL)中属于低精度类别。遥感基础模型Prithvi-EO在计算机视觉任务中表现出强迁移能力,但其视觉变压器(ViT)主干仅生成单一空间尺度特征,不适用于目标检测头所需的多尺度特征。现有方法通过缩放单步长令牌合成多尺度金字塔,牺牲了空间异质性;而全主干微调对基础模型计算成本过高。本文评估结合两种参数高效微调(PEFT)策略:低秩适配(LoRA)与混合PEFT,以及三种颈部设计:伪多尺度、Lite ViT-Adapter和全ViT-Adapter。最佳配置(Lite ViT-Adapter + 单阶段头)使用Diou损失,达到mAP@50 0.9479,表明中心感知定位对不规则休耕地块检测有效。在LoRA下,无适配器的一阶段检测比基线锚点方法提升6.42%,最佳配置提升25.70%。结果表明,轻量级空间先验融合与选择性主干解冻可使Prithvi-EO更有效地捕捉局部休耕模式,优于依赖重塑单步长令牌的方法。
原文摘要 · Abstract (English)
Understanding spatial distribution of fallow land is important for optimizing the food-water (FW) nexus, given fallowing's role in crop rotation and water conservation. Fallow is a low accuracy class in USDA Cropland Data Layer (CDL). Geospatial foundation model (GFM), Prithvi-EO has shown strong transferability across computer vision tasks. However, its Vision Transformer (ViT) backbone produces features at a single spatial scale that are ill-suited for the multi-scale features required by object detection heads. Existing approaches synthesise multi-scale pyramids through scaling of single stride tokens, sacrificing spatial heterogeneity, and full backbone fine-tuning is computationally prohibitive for GFMs. We evaluate a fallow detection pipeline combining two parameter-efficient fine tuning (PEFT) schemes: Low-Rank Adaptation (LoRA) and a hybrid PEFT, with three neck designs: pseudo multi-scale, Lite ViT-Adapter, and Full ViT-Adapter. Our best configuration, Lite ViT-Adapter with a one-stage head, achieves a mAP@50 of 0.9479 with the Diou loss, suggesting the effectiveness of center-aware localization for irregular fallow field detection. ViT-Adapter free one-stage detection under LoRA improves the adapter-free anchor-based approach by 6.42%, and the best configuration improves baseline adapter-free anchor-based approach by 25.70%. These results demonstrate that lightweight spatial prior fusion and selective backbone unfreezing enable Prithvi-EO to capture local fallow patterns more effectively, outperforming approaches that rely on reshaped single-stride ViT tokens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。