对比三类模型在农田病害识别中的表现,发现选型应根据实际场景决定。
AgriPath: A Systematic Exploration of Architectural Trade-offs for Crop Disease Classification
- 对比CNN、对比视觉语言模型和生成式视觉语言模型的性能差异
- 生成式模型抗域偏移最强,但存在自由文本生成失败问题
- 提出含16作物41病害的11.1万张图像新基准数据集
可靠作物病害检测需要模型在多样采集条件下保持稳定性能,但现有评估常局限于单一架构或实验室生成数据集。本文系统比较了三种用于细粒度作物病害分类的模型范式:卷积神经网络(CNN)、对比视觉语言模型(VLM)和生成式VLM。为控制领域效应分析,我们引入AgriPath-LF16,一个包含11.1万张图像的基准数据集,涵盖16种作物和41种病害,明确区分实验室与田间图像,并提供3万张平衡子集用于标准化训练与评估。所有模型在统一协议下,于全量、仅实验室、仅田间三种训练模式下进行训练与评估,采用宏平均F1分数与解析成功率(PSR)衡量生成可靠性(即输出可解析性)。结果表明:CNN在同域图像上准确率最高,但域偏移下显著退化;对比VLM提供稳健且参数高效的跨域性能;生成式VLM对分布变化最鲁棒,但存在因自由文本生成引发的额外失效模式。研究强调,架构选择应由部署环境决定,而非仅看总体性能。
原文摘要 · Abstract (English)
Reliable crop disease detection requires models that perform consistently across diverse acquisition conditions, yet existing evaluations often focus on single architectural families or lab-generated datasets. This work presents a systematic empirical comparison of three model paradigms for fine-grained crop disease classification: Convolutional Neural Networks (CNNs), contrastive Vision-Language Models (VLMs), and generative VLMs. To enable controlled analysis of domain effects, we introduce AgriPath-LF16, a benchmark of 111k images spanning 16 crops and 41 diseases with explicit separation between laboratory and field imagery, alongside a balanced 30k subset for standardised training and evaluation. We train and evaluate all models under unified protocols across full, lab-only, and field-only training regimes using macro-F1 and Parse Success Rate (PSR) to account for generative reliability (i.e., output parsability measured via PSR). The results reveal distinct performance profiles: CNNs achieve the highest accuracy on in-domain imagery but exhibit pronounced degradation under domain shift; contrastive VLMs provide a robust and parameter-efficient alternative with competitive cross-domain performance; generative VLMs demonstrate the strongest resilience to distributional variation, albeit with additional failure modes stemming from free-text generation. These findings highlight that architectural choice should be guided by deployment context rather than aggregate performance alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。