用260万张田间图像训练的扩散模型,提升作物结构感知能力。
SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- 基于像素空间扩散变压器,从无标注图像中学习结构保持的表征
- 在器官分割、杂草识别等任务上显著优于主流预训练模型
- 适合需要低标注成本的农业视觉研究者使用
基于图像的植物表型分析依赖对作物密集结构的精准理解,但跨物种、器官和生长阶段的像素级标注成本高昂。通用视觉基础模型虽能提升标签效率,但其网络规模预训练目标在农业图像上迁移效果弱,因农业图像语义常由重复纹理场景中的精细器官几何决定。本文提出SPROUT,一种用于多作物植物表型分析的扩散基础模型。SPROUT利用260万张无标注田间图像(MCD-2.6M)进行训练,采用像素空间扩散变压器,并通过去噪时间步上的无标签有效秩准则筛选可迁移特征。该设计将预训练重点从作物不变性转向结构保持去噪,使表示更契合密集表型任务。我们在器官分割、作物-杂草解析、深度估计和计数等任务上评估SPROUT,结果表明其持续优于强基线模型,尤其在密集结构预测上提升显著,且相比通用与作物专用基础模型具有更优的标注与计算效率。源代码与MCD-2.6M数据集已公开。
原文摘要 · Abstract (English)
Image-based plant phenotyping depends on dense structural understanding of crops, yet pixel-level annotation remains expensive across species, organs, growth stages, and field conditions. General-purpose vision foundation models offer a natural route to label efficiency, but their web-scale pretraining objectives transfer weakly to agricultural imagery, where semantics are often determined by fine organ geometry inside repetitive, texture-dominated scenes. We introduce SPROUT, a diffusion foundation model for multi-crop plant phenotyping. SPROUT learns from 2.6 million unlabeled open-field images (MCD-2.6M) using a pixel-space Diffusion Transformer, and selects transferable features with a label-free effective-rank criterion over denoising timesteps. This design shifts pretraining from crop-based invariance to structure-preserving denoising, making the representation better aligned with dense phenotyping tasks. We evaluate SPROUT across dense phenotyping tasks, including organ segmentation, crop-weed parsing, depth estimation, and counting. SPROUT consistently improves over strong web-pretrained baselines, with the largest gains on dense structural prediction, and shows favorable label and compute efficiency compared with general-purpose and crop-specific foundation models. The source code and MCD-2.6M dataset are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。