arXiv:2604.09907cs.CVcs.AI2026-04

构建植物表型多模态推理基准,提升农业大模型精准分析能力

From UAV Imagery to Agronomic Reasoning: A Multimodal LLM Benchmark for Plant Phenotyping

论文配图:From UAV Imagery to Agronomic Reasoning: A Multimodal LLM Benchmark for Plant Phenotyping
图 1 · 摘自论文原文
  • 设计面向大豆棉花的多模态推理评测集,支持视觉与农学知识结合分析
  • 11个顶尖视觉语言模型测试中,微调后最高准确率达78%
  • 揭示模型规模增长受限、跨作物泛化不均等关键挑战

为提升作物遗传改良效率,高通量、高效且全面的表型分析至关重要。传统人工方式正被多模态基础模型(尤其是视觉-语言模型,VLMs)推动的自动化分析所取代。然而,植物科学对领域知识、细粒度视觉理解及复杂生物农学推理要求极高,仍具挑战性。为此,我们构建了PlantXpert——一个基于证据的多模态推理基准,专用于大豆和棉花表型分析。该基准提供结构化、可复现的农学适配框架,支持基线模型与领域微调模型的对照评估。数据集包含385张数字图像和超过3000个评测样本,覆盖病害、虫害、杂草管理与产量等核心领域。可评估视觉专业性、定量推理及多步农学推理等能力。共评估11个前沿VLMs,结果表明任务微调显著提升精度,如Qwen3-VL-4B与Qwen3-VL-30B最高达78%。但模型规模收益递减,跨作物泛化能力不均衡,定量与生物学依据推理仍存重大困难。这表明PlantXpert可作为评估证据驱动农学推理的基础,并推动植物科学多模态模型发展。

原文摘要 · Abstract (English)

To improve crop genetics, high-throughput, effective and comprehensive phenotyping is a critical prerequisite. While such tasks were traditionally performed manually, recent advances in multimodal foundation models, especially in vision-language models (VLMs), have enabled more automated and robust phenotypic analysis. However, plant science remains a particularly challenging domain for foundation models because it requires domain-specific knowledge, fine-grained visual interpretation, and complex biological and agronomic reasoning. To address this gap, we develop PlantXpert, an evidence-grounded multimodal reasoning benchmark for soybean and cotton phenotyping. Our benchmark provides a structured and reproducible framework for agronomic adaptation of VLMs, and enables controlled comparison between base models and their domain-adapted counterparts. We constructed a dataset comprising 385 digital images and more than 3,000 benchmark samples spanning key plant science domains including disease, pest control, weed management, and yield. The benchmark can assess diverse capabilities including visual expertise, quantitative reasoning, and multi-step agronomic reasoning. A total of 11 state-of-the-art VLMs were evaluated. The results indicate that task-specific fine-tuning leads to substantial improvement in accuracy, with models such as Qwen3-VL-4B and Qwen3-VL-30B achieving up to 78%. At the same time, gains from model scaling diminish beyond a certain capacity, generalization across soybean and cotton remains uneven, and quantitative as well as biologically grounded reasoning continue to pose substantial challenges. These findings suggest that PlantXpert can serve as a foundation for assessing evidence-grounded agronomic reasoning and for advancing multimodal model development in plant science.

多模态表型分析农业大模型视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。