让AI像植物病理学家一样逐步提问诊断,提升诊断准确性。
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry

- 构建多步推理框架,基于视觉线索和诊断意图生成问题链。
- 模型在复杂诊断中准确率低,但结构化提问可减少幻觉并提效。
- 适合研究医疗诊断、多步推理与可解释AI的学者参考。
视觉评估通常依赖多步过程。在植物病理学中,专家通过结构化、基于证据的动态提问分析叶面图像,识别视觉线索,推断诊断意图,并根据物种、症状和严重程度提出针对性问题。这一流程对准确诊断和治疗至关重要。然而,当前视觉语言模型多采用单轮问答评估。为此,我们提出PlantInquiryVQA基准,用于研究植物诊断中的多步、意图驱动视觉推理。我们建立了一个包含24,950张专家标注植物图像和138,068个问答对的数据集,每个问答对均标注视觉定位、严重程度及领域特定推理模板。在顶级多模态大模型上的评估显示,尽管模型能较好描述症状,但在安全临床推理和准确诊断上表现不佳。重要的是,结构化问题引导显著提升了诊断正确率,减少了幻觉,并提高了推理效率。我们希望PlantInquiryVQA成为推动诊断智能体向专家级推理演进的基础基准。
原文摘要 · Abstract (English)
Vision evaluations are typically done through multi-step processes. In most contemporary fields, experts analyze images using structured, evidence-based adaptive questioning. In plant pathology, botanists inspect leaf images, identify visual cues, infer diagnostic intent, and probe further with targeted questions that adapt to species, symptoms, and severity. This structured probing is crucial for accurate disease diagnosis and treatment formulation. Yet current vision-language models are evaluated on single-turn question answering. To address this gap, we introduce PlantInquiryVQA, a benchmark for studying multi-step, intent-driven visual reasoning in botanical diagnosis. We formalize a Chain of Inquiry framework modeling diagnostic trajectories as ordered question-answer sequences conditioned on grounded visual cues and explicit epistemic intent. We release a dataset of 24,950 expert-curated plant images and 138,068 question-answer pairs annotated with visual grounding, severity labels, and domain-specific reasoning templates. Evaluations on top-tier Multimodal Large Language Models reveal that while they describe visual symptoms adequately, they struggle with safe clinical reasoning and accurate diagnosis. Importantly, structured question-guided inquiry significantly improves diagnostic correctness, reduces hallucination, and increases reasoning efficiency. We hope PlantInquiryVQA serves as a foundational benchmark in advancing research to train diagnostic agents to reason like expert botanists rather than static classifiers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。