arXiv:2605.01283cs.CVcs.AI2026-05

构建新基线模型,提升植物叶片病害识别效率与泛化能力

Developing a Strong Pre-Trained Base Model for Plant Leaf Disease Classification

  • 基于DenseNet201架构,融合多数据集与增强策略构建新训练集
  • 在两个新数据集上实现比基准模型更快、更稳定、更少数据的迁移学习效果
  • 为农业病害检测提供高效、鲁棒的预训练基础模型,适合资源受限场景

植物、农作物及其产量关乎人类生存,但病害与虫害每年造成巨大损失。早期发现并及时处理对遏制传播至关重要。传统人工巡查耗时费力,机器学习方法因此被引入,其中卷积神经网络(CNN)能自动提取图像特征,表现优异。然而,模型性能高度依赖数据集质量。尽管数据集重要,公开可用资源仍难以满足训练高性能模型的需求。为此,本研究梳理了现有公开数据集,评估其适用性,并通过增强策略研究优化数据构建。基于此,构建了一个新数据集,用于训练基于DenseNet201的新型基础模型。该模型在新数据集上超越基线,且在另一次域内迁移学习实验中表现更优,具备更快收敛、更强鲁棒性、更低数据需求等优势,显著缓解了该领域长期存在的训练效率与泛化难题。

原文摘要 · Abstract (English)

Plants, crops and their yields are essential to our very existence, but diseases and pests cause large losses every year. As such it is vital to ensure that diseases can be spotted early and treated accordingly and stopping the spread while still possible. Manual and traditional methods require personal to walk through the field and check for symptoms 'by hand'. This is very laborious and very time consuming, so ML methods have been applied as a result and they have garnered promising results. CNN models are especially efficient as they can automatically extract features from images without any manual feature construction before then feeding the features to a classifier. Datasets are largely influential to the final performance of the model. Despite the importance that datasets pose to the field, there still seems to be somewhat of a discrepancy between what is publicly available for use and what would be required to sufficiently train fully capable models. To overcome these shortcomings, as part of this thesis open datasets for the field of plant leaf disease classification have been identified as well as models that can be trained on them and extensive benchmarks have been carried out to identify their suitability. Then a new dataset was constructed based on those findings as well as on the findings of a augmentation applicability study, which will be used to train a new Base Model based on the DenseNet201 architecture, which managed to outperform the baseline model on said new dataset as well as outperforming it on plant leaf disease classification domain specific Transfer-Learning experiments on another new dataset. This new model manages to train models through Transfer-Learning (TL) faster, more robust, more stable, and with less data than general model would, overcoming a large number of issues that the field still suffers from.

病害识别迁移学习植物视觉预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。