arXiv:2607.16492cs.CV2026-07

用强化学习让AI精准识别杂草,还能解释原因。

WeedExpert-R1: Incentivizing Botanical Reasoning in MLLMs with Reinforcement Learning for Precision Weed Grounding

论文配图:WeedExpert-R1: Incentivizing Botanical Reasoning in MLLMs with Reinforcement Learning for Precision Weed Grounding
图 1 · 摘自论文原文
  • 用植物特征库+人类审核生成推理数据,训练模型理解杂草
  • 在37种杂草上准确率达89.3%,实例召回率87.8%
  • 无需重新训练就能识别新杂草,适合多地区农田部署

精准除草需要物种级识别和实例级定位。传统目标检测器采用封闭词表,限制了跨区域应用,且无法在复杂农田场景中解释预测结果。多模态大语言模型(MLLM)具备视觉定位与推理能力,但缺乏足够的植物学知识易导致细粒度杂草识别中的幻觉。本研究提出WeedExpert-R1,通过可验证奖励机制学习视觉引导的植物学推理。利用人工校准的植物性状词典与审计-合成型LLM工作流,构建推理数据用于监督微调;再采用分组相对策略优化,奖励格式、准确性、实例数量和回答长度。在六个数据集共37种杂草上,WeedExpert-R1-4B在0.5交并比阈值下达到75.82%精确集精度,89.30%精确率,87.81%召回率。其性能优于商用模型(如GPT-5.4、Gemini-3.1-Pro)及更大开源模型(如Qwen3-VL-30B-Instruct、Gemma-4-31B-it)。对未见物种的测试进一步证明其开放词汇能力,具备无需重训即可在不同地区和作物中部署的潜力。

原文摘要 · Abstract (English)

Precision weed control requires species-level identification and instance-level localization. However, conventional object detectors use a closed vocabulary, limiting their deployment across regions, and cannot explain their predictions in complex agricultural scenes. Multimodal large language models (MLLMs) offer visual grounding and reasoning capabilities, but insufficient botanical knowledge can cause hallucinations in fine-grained weed identification. This study introduces WeedExpert-R1, a multimodal model that learns visually grounded botanical reasoning through verifiable rewards. A domain-specific Chain-of-Thought synthesis pipeline combines a human-curated botanical trait dictionary with an Auditor-Synthesizer LLM workflow to generate reasoning data for supervised fine-tuning. Group Relative Policy Optimization is then applied with rewards for format, accuracy, instance count, and response length. Across 37 weed species from six datasets, WeedExpert-R1-4B achieved 75.82 percent exact-set precision at an IoU threshold of 0.5, 89.30 percent precision, and 87.81 percent recall. It outperformed proprietary models, including GPT-5.4 and Gemini-3.1-Pro, and larger open-source models, including Qwen3-VL-30B-Instruct and Gemma-4-31B-it. Results on unseen species further demonstrate its open-vocabulary capability and potential for deployment across diverse regions and crops without retraining.

农业AI视觉推理多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。