arXiv:2603.16934cs.CVcs.AI2026-03被引 1

用可验证的农业知识生成数据,训练出更可靠的农业多模态模型。

AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding

  • 通过视觉描述+科学文献检索自动生成带验证的知识数据
  • 构建超3000类、60万+图文问答的农业基准数据集
  • 适合农业科研与智慧农业从业者使用

当前农业多模态大模型的发展受限于两大瓶颈:缺乏大规模农业数据集,且现有模型缺乏经验证的领域专业知识以跨分类体系推理。为此,我们提出视觉到可验证知识(V2VK)管道,利用生成式AI结合网络增强的科学文献检索,自动构建AgriMM基准数据集,有效避免生物幻觉,确保训练数据基于权威植物病理学文献。AgriMM包含超过3,000个农业类别和607,000多个视觉-问答对,覆盖细粒度植物识别、病害症状判断、作物计数与成熟度评估等任务。基于此可验证数据,我们提出AgriChat,一个专注于农业领域的多模态大模型,具备数千类农业知识,并能提供详细解释的农业评估。在多种任务、数据集与评估条件下广泛测试表明,该模型在开放源代码模型中表现最优,验证了保留视觉细节并结合网络验证知识是构建可靠可信农业AI的有效路径。代码与数据集已公开于https://github.com/boudiafA/AgriChat。

原文摘要 · Abstract (English)

The deployment of Multimodal Large Language Models (MLLMs) in agriculture is currently stalled by a critical trade-off: the existing literature lacks the large-scale agricultural datasets required for robust model development and evaluation, while current state-of-the-art models lack the verified domain expertise necessary to reason across diverse taxonomies. To address these challenges, we propose the Vision-to-Verified-Knowledge (V2VK) pipeline, a novel generative AI-driven annotation framework that integrates visual captioning with web-augmented scientific retrieval to autonomously generate the AgriMM benchmark, effectively eliminating biological hallucinations by grounding training data in verified phytopathological literature. The AgriMM benchmark contains over 3,000 agricultural classes and more than 607k VQAs spanning multiple tasks, including fine-grained plant species identification, plant disease symptom recognition, crop counting, and ripeness assessment. Leveraging this verifiable data, we present AgriChat, a specialized MLLM that presents broad knowledge across thousands of agricultural classes and provides detailed agricultural assessments with extensive explanations. Extensive evaluation across diverse tasks, datasets, and evaluation conditions reveals both the capabilities and limitations of current agricultural MLLMs, while demonstrating AgriChat's superior performance over other open-source models, including internal and external benchmarks. The results validate that preserving visual detail combined with web-verified knowledge constitutes a reliable pathway toward robust and trustworthy agricultural AI. The code and dataset are publicly available at https://github.com/boudiafA/AgriChat .

农业AI多模态模型知识验证数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。