arXiv:2503.06973cs.CVcs.AI2025-03ECCV被引 38

构建首个作物病害多模态数据集,助力AI精准诊断与农技问答。

A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis

  • 融合13.7万张病害图像与百万级图文问答对,支持多模态理解。
  • 采用低秩适配微调策略,显著提升视觉-语言模型诊断准确率。
  • 适合农业AI研究者、智能农技系统开发者使用。

尽管对话式生成AI在提升农业决策方面展现出巨大潜力,但其应用仍以文本交互为主。多模态对话AI借助海量跨源图文数据取得显著进展,但在作物病害诊断等农业场景中仍处于探索阶段。本文提出作物病害领域多模态(CDDM)数据集,包含137,000张作物病害图像及100万条涵盖病害识别与管理实践的问答对,推动农业多模态学习发展。通过整合视觉与文本信息,该数据集支持开发高精度问答系统,为农民与农业从业者提供实用建议。我们采用一种新颖的微调策略,联合适配视觉编码器、适配器与语言模型,基于低秩适应(LoRA)实现高效优化。实验表明该方法在作物病害诊断任务中性能显著提升。本工作贡献包括数据集、微调策略与基准测试,旨在弥合先进AI技术与农业应用间的鸿沟。数据集已开源:https://github.com/UnicomAI/UnicomBenchmark/tree/main/CDDMBench。

原文摘要 · Abstract (English)

While conversational generative AI has shown considerable potential in enhancing decision-making for agricultural professionals, its exploration has predominantly been anchored in text-based interactions. The evolution of multimodal conversational AI, leveraging vast amounts of image-text data from diverse sources, marks a significant stride forward. However, the application of such advanced vision-language models in the agricultural domain, particularly for crop disease diagnosis, remains underexplored. In this work, we present the crop disease domain multimodal (CDDM) dataset, a pioneering resource designed to advance the field of agricultural research through the application of multimodal learning techniques. The dataset comprises 137,000 images of various crop diseases, accompanied by 1 million question-answer pairs that span a broad spectrum of agricultural knowledge, from disease identification to management practices. By integrating visual and textual data, CDDM facilitates the development of sophisticated question-answering systems capable of providing precise, useful advice to farmers and agricultural professionals. We demonstrate the utility of the dataset by finetuning state-of-the-art multimodal models, showcasing significant improvements in crop disease diagnosis. Specifically, we employed a novel finetuning strategy that utilizes low-rank adaptation (LoRA) to finetune the visual encoder, adapter and language model simultaneously. Our contributions include not only the dataset but also a finetuning strategy and a benchmark to stimulate further research in agricultural technology, aiming to bridge the gap between advanced AI techniques and practical agricultural applications. The dataset is available at https: //github.com/UnicomAI/UnicomBenchmark/tree/main/CDDMBench.

作物病害多模态AI助农数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。