打造农业视觉语言理解新框架,含海量数据、专用模型与评测体系。
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
- 构建100万图像-标题对的农业多模态语料库,由多智能体生成。
- 模型在农业评测中超越通用大模型,且不损失纯文本能力。
- 开源完整资源,适合农业AI研究者与低资源场景应用。
尽管多模态大模型发展迅速,农业应用仍受限于领域专用模型、精心构建的视觉语言语料库及严格评估的缺乏。为此,我们提出AgriGPT-VL套件,一个统一的农业多模态框架。第一,引入Agri-3M-VL,目前已知最大的农业视觉语言语料库,通过可扩展的多智能体数据生成器构建,包含100万图像-标题对、200万图像支撑的VQA对、5万专家级VQA实例和1.5万GRPO强化学习样本。第二,开发AgriGPT-VL,一种基于渐进式课程学习的农业专用多模态模型,经文本定位、多模态浅层/深层对齐及GRPO精炼训练,具备强多模态推理能力并保留纯文本性能。第三,建立AgriBench-VL-4K,一个紧凑但具挑战性的评测套件,含开放式和图像依赖问题,搭配多指标评估与LLM-as-a-judge框架。实验显示,AgriGPT-VL在AgriBench-VL-4K上优于主流通用多模态模型,且在LLM-as-a-judge评估中胜率更高;同时在文本类评测AgriBench-13K上保持竞争力,无明显语言能力下降。消融实验进一步验证对齐与GRPO精炼阶段的持续增益。所有资源将开源,以支持可复现研究及低资源农业场景部署。
原文摘要 · Abstract (English)
Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the scarcity of domain-tailored models, curated vision-language corpora, and rigorous evaluation. To address these challenges, we present the AgriGPT-VL Suite, a unified multimodal framework for agriculture. Our contributions are threefold. First, we introduce Agri-3M-VL, the largest vision-language corpus for agriculture to our knowledge, curated by a scalable multi-agent data generator; it comprises 1M image-caption pairs, 2M image-grounded VQA pairs, 50K expert-level VQA instances, and 15K GRPO reinforcement learning samples. Second, we develop AgriGPT-VL, an agriculture-specialized vision-language model trained via a progressive curriculum of textual grounding, multimodal shallow/deep alignment, and GRPO refinement. This method achieves strong multimodal reasoning while preserving text-only capability. Third, we establish AgriBench-VL-4K, a compact yet challenging evaluation suite with open-ended and image-grounded questions, paired with multi-metric evaluation and an LLM-as-a-judge framework. Experiments show that AgriGPT-VL outperforms leading general-purpose VLMs on AgriBench-VL-4K, achieving higher pairwise win rates in the LLM-as-a-judge evaluation. Meanwhile, it remains competitive on the text-only AgriBench-13K with no noticeable degradation of language ability. Ablation studies further confirm consistent gains from our alignment and GRPO refinement stages. We will open source all of the resources to support reproducible research and deployment in low-resource agricultural settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。