arXiv:2504.20419cs.CV2025-04被引 14

用大模型+卷积网络自动识别植物病害,效果优于传统方法。

Plant Disease Detection through Multimodal Large Language Models and Convolutional Neural Networks

  • 结合GPT-4o与CNN,通过微调提升病叶分类能力。
  • 微调后对苹果叶识别准确率达98.12%,超ResNet-50的96.88%。
  • 适合缺乏标注数据的农业场景,降低对高分辨率设备依赖。

农业自动化在作物监测与病害管理中至关重要,尤其依赖早期检测系统。本研究探究将多模态大语言模型(如GPT-4o)与卷积神经网络(CNN)结合,用于基于叶片图像的植物病害自动分类。基于PlantVillage数据集,系统评估了零样本、少样本及渐进式微调场景下的模型表现。对比GPT-4o与广泛使用的ResNet-50,在三种分辨率(100、150、256像素)和两种植物(苹果、玉米)上进行分析。结果显示,微调后的GPT-4o性能略优,对苹果叶图像分类准确率最高达98.12%,高于ResNet-50的96.88%,且泛化能力更强,训练损失接近零。但零样本表现显著偏低,表明需少量训练。跨分辨率与跨植物泛化测试揭示了模型在新场景中的适应性与局限性。结果表明,融合多模态大模型可提升病害检测系统的可扩展性与智能水平,减少对大规模标注数据和高分辨率传感器的依赖。

原文摘要 · Abstract (English)

Automation in agriculture plays a vital role in addressing challenges related to crop monitoring and disease management, particularly through early detection systems. This study investigates the effectiveness of combining multimodal Large Language Models (LLMs), specifically GPT-4o, with Convolutional Neural Networks (CNNs) for automated plant disease classification using leaf imagery. Leveraging the PlantVillage dataset, we systematically evaluate model performance across zero-shot, few-shot, and progressive fine-tuning scenarios. A comparative analysis between GPT-4o and the widely used ResNet-50 model was conducted across three resolutions (100, 150, and 256 pixels) and two plant species (apple and corn). Results indicate that fine-tuned GPT-4o models achieved slightly better performance compared to the performance of ResNet-50, achieving up to 98.12% classification accuracy on apple leaf images, compared to 96.88% achieved by ResNet-50, with improved generalization and near-zero training loss. However, zero-shot performance of GPT-4o was significantly lower, underscoring the need for minimal training. Additional evaluations on cross-resolution and cross-plant generalization revealed the models' adaptability and limitations when applied to new domains. The findings highlight the promise of integrating multimodal LLMs into automated disease detection pipelines, enhancing the scalability and intelligence of precision agriculture systems while reducing the dependence on large, labeled datasets and high-resolution sensor infrastructure. Large Language Models, Vision Language Models, LLMs and CNNs, Disease Detection with Vision Language Models, VLMs

病害检测大模型农业AI多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。