arXiv:2506.20100cs.LGcs.AI2025-06NeurIPS被引 12

构建农业专家对话多模态推理基准,支持真实场景下的复杂决策评估

MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations

  • 基于3.5万条真实用户-专家交互数据构建多模态问答集
  • 涵盖7000+生物实体,覆盖作物、病害与虫害的开放世界诊断
  • 强调模型在模糊输入下主动澄清与长文本生成能力,适合农业AI研究

我们提出MIRAGE,一个面向农业领域专家咨询对话中多模态信息寻求与推理的新基准。该基准通过整合自然用户提问、专家撰写回复及图像上下文,真实还原专家咨询的复杂性,可用于评估模型在具身推理、澄清策略与长篇生成方面的能力。基于超过35,000条真实用户-专家交互数据,经多阶段精心筛选,MIRAGE覆盖多样化的作物健康、病虫害诊断与作物管理场景。其包含逾7,000个独特生物实体,涵盖植物种类、害虫与病害,是目前最具有分类多样性的视觉语言模型基准之一,且扎根于真实世界。与依赖明确输入和封闭类别体系的现有基准不同,MIRAGE采用不明确、富含上下文的开放世界设定,要求模型推断潜在知识缺口,处理罕见实体,或主动引导对话进程或作出响应。

原文摘要 · Abstract (English)

We introduce MIRAGE, a new benchmark for multimodal expert-level reasoning and decision-making in consultative interaction settings. Designed for the agriculture domain, MIRAGE captures the full complexity of expert consultations by combining natural user queries, expert-authored responses, and image-based context, offering a high-fidelity benchmark for evaluating models on grounded reasoning, clarification strategies, and long-form generation in a real-world, knowledge-intensive domain. Grounded in over 35,000 real user-expert interactions and curated through a carefully designed multi-step pipeline, MIRAGE spans diverse crop health, pest diagnosis, and crop management scenarios. The benchmark includes more than 7,000 unique biological entities, covering plant species, pests, and diseases, making it one of the most taxonomically diverse benchmarks available for vision-language models, grounded in the real world. Unlike existing benchmarks that rely on well-specified user inputs and closed-set taxonomies, MIRAGE features underspecified, context-rich scenarios with open-world settings, requiring models to infer latent knowledge gaps, handle rare entities, and either proactively guide the interaction or respond. Project Page: https://mirage-benchmark.github.io

多模态推理农业AI对话系统视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。