arXiv:2506.06084cs.CV2025-06被引 5

构建小麦管理专用视觉语言数据集,提升模型精准决策能力。

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management

  • 设计三层数据集,覆盖形态预训练、量化测量与管理指令微调
  • 微调后模型在病害诊断和生育期识别上准确率达79.2%和84.6%
  • 适合农业智能决策、遥感分析及多模态模型研究者使用

小麦管理策略对产量至关重要。传统依赖人工专家巡检的方法成本高、主观性强且难规模化。近期视觉语言模型(VLMs)为实现可扩展的数据驱动管理支持提供了新路径。然而,因缺乏领域知识,直接应用VLM于小麦管理任务时,量化与推理能力差,常产生模糊甚至误导性建议。为此,我们提出WisWheat,一个三层次小麦专用数据集:(1) 包含47,871张图像-文本对的基础预训练数据集,用于粗粒度适配模型至小麦形态;(2) 7,263个VQA风格的图文问答对,支持定量性状测量任务;(3) 4,888个指令微调样本,聚焦不同生育期的生物与非生物胁迫诊断及管理方案生成。大量实验表明,将开源VLM(如Qwen2.5 7B)在该数据集上微调后性能显著提升。特别是,基于小麦指令数据集微调的Qwen2.5 VL 7B在胁迫识别与生育期对话任务中分别达到79.2%和84.6%准确率,优于通用商业模型GPT-4o 11.9%和34.6%。

原文摘要 · Abstract (English)

Wheat management strategies play a critical role in determining yield. Traditional management decisions often rely on labour-intensive expert inspections, which are expensive, subjective and difficult to scale. Recently, Vision-Language Models (VLMs) have emerged as a promising solution to enable scalable, data-driven management support. However, due to a lack of domain-specific knowledge, directly applying VLMs to wheat management tasks results in poor quantification and reasoning capabilities, ultimately producing vague or even misleading management recommendations. In response, we propose WisWheat, a wheat-specific dataset with a three-layered design to enhance VLM performance on wheat management tasks: (1) a foundational pretraining dataset of 47,871 image-caption pairs for coarsely adapting VLMs to wheat morphology; (2) a quantitative dataset comprising 7,263 VQA-style image-question-answer triplets for quantitative trait measuring tasks; and (3) an Instruction Fine-tuning dataset with 4,888 samples targeting biotic and abiotic stress diagnosis and management plan for different phenological stages. Extensive experimental results demonstrate that fine-tuning open-source VLMs (e.g., Qwen2.5 7B) on our dataset leads to significant performance improvements. Specifically, the Qwen2.5 VL 7B fine-tuned on our wheat instruction dataset achieves accuracy scores of 79.2% and 84.6% on wheat stress and growth stage conversation tasks respectively, surpassing even general-purpose commercial models such as GPT-4o by a margin of 11.9% and 34.6%.

视觉语言小麦管理数据集农业AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。