arXiv:2509.06907cs.CV2025-09被引 6

为小麦视觉建模构建了全球多源数据预训练模型,提升田间识别可靠性。

FoMo4Wheat: Toward reliable crop vision foundation models with globally curated data

  • 基于250万张小麦图像自监督预训练,覆盖30个全球站点
  • 在10项田间任务中均优于通用模型,跨物种泛化能力强
  • 适合农业视觉研究者和数字育种团队使用

以视觉驱动的田间监测是数字农业的核心,但基于通用领域预训练模型的系统常因植株冠层结构细微差异与田间环境波动而难以泛化。本文提出FoMo4Wheat,首个基于自监督学习在全球最大最多样小麦图像数据集ImAg4Wheat上预训练的作物领域视觉基础模型。该数据集包含十年间在30个全球站点采集的250万张高分辨率图像,覆盖超2000个基因型和500种以上环境条件。小麦专项预训练使模型表征对小麦具有鲁棒性,并可迁移至其他作物与杂草识别任务。在10项冠层及器官级田间视觉任务中,FoMo4Wheat始终优于通用数据集预训练的最先进模型。结果证明作物专用基础模型对可靠田间感知的价值,并为实现跨物种、跨任务的通用作物基础模型指明路径。模型与数据集已公开:https://github.com/PheniX-Lab/FoMo4Wheat 及 https://huggingface.co/PheniX-Lab/FoMo4Wheat。演示网站:https://fomo4wheat.phenix-lab.com/

原文摘要 · Abstract (English)

Vision-driven field monitoring is central to digital agriculture, yet models built on general-domain pretrained backbones often fail to generalize across tasks, owing to the interaction of fine, variable canopy structures with fluctuating field conditions. We present FoMo4Wheat, one of the first crop-domain vision foundation model pretrained with self-supervision on ImAg4Wheat, the largest and most diverse wheat image dataset to date (2.5 million high-resolution images collected over a decade at 30 global sites, spanning >2,000 genotypes and >500 environmental conditions). This wheat-specific pretraining yields representations that are robust for wheat and transferable to other crops and weeds. Across ten in-field vision tasks at canopy and organ levels, FoMo4Wheat models consistently outperform state-of-the-art models pretrained on general-domain dataset. These results demonstrate the value of crop-specific foundation models for reliable in-field perception and chart a path toward a universal crop foundation model with cross-species and cross-task capabilities. FoMo4Wheat models and the ImAg4Wheat dataset are publicly available online: https://github.com/PheniX-Lab/FoMo4Wheat and https://huggingface.co/PheniX-Lab/FoMo4Wheat. The demonstration website is: https://fomo4wheat.phenix-lab.com/.

作物视觉基础模型数字农业自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。