arXiv:2608.07984cs.CV2026-08中稿 · the Computer Visio…

用4万张农田图像高效训练农业视觉模型,节省9倍参数量。

AgriField-40K: Adapting Vision Models to Agriculture With Efficient Continual Pretraining

论文配图:AgriField-40K: Adapting Vision Models to Agriculture With Efficient Continual Pretraining
图 1 · 摘自论文原文
  • 仅训练轻量适配器,实现参数高效持续预训练
  • 在多任务上表现优于全量微调,参数减少9倍
  • 适合农业视觉研究者快速迁移学习

田间农业计算机视觉对精准农业至关重要,但依赖昂贵标注和大型模型的高成本适配。我们提出AgriField-40K,一个从17个公开资源收集的田间数据集,涵盖多种作物、杂草及田间环境。基于此,我们构建AgriMAE,一种参数高效的持续预训练基线,通过仅训练轻量适配器来适应自然图像预训练的掩码自编码器。我们进一步探索语义特征重建作为替代预训练目标,并评估跨任务迁移效果。AgriMAE在下游任务中持续提升性能,使用最多减少9倍可训练参数的情况下,仍可匹配甚至超越全量微调表现,证明AgriField-40K是农业视觉持续预训练的实用资源。

原文摘要 · Abstract (English)

Field-based agricultural computer vision is important for precision agriculture, yet it largely depends on expensive annotations and costly adaptation of large pretrained models. We introduce AgriField-40K, a field-centric dataset curated from 17 public resources and covering diverse crops, weeds, and field conditions. Building on this, we present AgriMAE, a parameter-efficient continual pretraining baseline that adapts a masked autoencoder pretrained on natural images by training only lightweight adapters. We further explore semantic feature reconstruction as an alternative pretraining objective and evaluate transfer across multiple tasks. AgriMAE consistently improves downstream performance and can match or even outperform full fine-tuning while using up to $9\times$ fewer trainable parameters, showing that AgriField-40K is a practical resource for continual pretraining in agricultural vision. Project page: https://dtu-pas.github.io/agrifield40k/

农业视觉持续预训练参数效率掩码自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。