用多模态大模型实现可交互的农作物病害智能诊断
AgriDoctor: A Multimodal Intelligent Assistant for Agriculture
- 构建模块化框架,融合视觉、语言与农业知识库
- 在40万张图像上训练,诊断准确率超越现有模型
- 适合农业科研、智慧农场和农业科技开发者使用
精准的作物病害诊断对可持续农业和全球粮食安全至关重要。现有方法主要依赖图像分类器等单模态模型,难以融入领域知识,且缺乏语言交互能力。尽管大语言模型(LLMs)和大视觉语言模型(LVLMs)为多模态推理带来新可能,但其在农业场景中的表现受限于缺乏专用数据集和领域适配不足。本文提出AgriDoctor,一个面向作物病害智能诊断与农业知识交互的模块化、可扩展多模态框架。作为首个将代理式多模态推理引入农业领域的尝试,AgriDoctor集成路由、分类、检测、知识检索和大语言模型五大组件。为支持有效训练与评估,我们构建了AgriMM基准,包含40万张标注病害图像、831条专家整理的知识条目以及30万条双语指令用于意图驱动的工具选择。大量实验表明,基于AgriMM训练的AgriDoctor在细粒度农业任务上显著优于当前最优的LVLMs,为智能可持续农业应用树立了新范式。
原文摘要 · Abstract (English)
Accurate crop disease diagnosis is essential for sustainable agriculture and global food security. Existing methods, which primarily rely on unimodal models such as image-based classifiers and object detectors, are limited in their ability to incorporate domain-specific agricultural knowledge and lack support for interactive, language-based understanding. Recent advances in large language models (LLMs) and large vision-language models (LVLMs) have opened new avenues for multimodal reasoning. However, their performance in agricultural contexts remains limited due to the absence of specialized datasets and insufficient domain adaptation. In this work, we propose AgriDoctor, a modular and extensible multimodal framework designed for intelligent crop disease diagnosis and agricultural knowledge interaction. As a pioneering effort to introduce agent-based multimodal reasoning into the agricultural domain, AgriDoctor offers a novel paradigm for building interactive and domain-adaptive crop health solutions. It integrates five core components: a router, classifier, detector, knowledge retriever and LLMs. To facilitate effective training and evaluation, we construct AgriMM, a comprehensive benchmark comprising 400000 annotated disease images, 831 expert-curated knowledge entries, and 300000 bilingual prompts for intent-driven tool selection. Extensive experiments demonstrate that AgriDoctor, trained on AgriMM, significantly outperforms state-of-the-art LVLMs on fine-grained agricultural tasks, establishing a new paradigm for intelligent and sustainable farming applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。