arXiv:2605.27595cs.CVcs.AI2026-05

研究多模态大模型在农业图像中生成和解读时的幻觉问题。

Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks

论文配图:Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks
图 1 · 摘自论文原文
  • 分析模型在图像理解与文本生成中的生物和环境一致性错误。
  • 零样本准确率63%~75%,少样本提升至86.8%,仍存误检漏检。
  • 顶尖模型生成91%不合理的农业场景,暴露本质缺陷。

大型语言模型(LLMs)正被快速应用于农业成像任务,涵盖作物解读与合成田间图像生成。然而,这些模型常产生看似自信却违背生物学或环境现实的幻觉输出,可能导致错误的农学判断。本研究从两个互补方向考察此类幻觉:图像到文本(如用模型描述作物或田块图像中的生物/非生物胁迫状况),以及文本到图像(根据描述生成合成农业场景)。我们评估了生物不一致、上下文错误与农学不合理性等错误,基于领域知识标准在多种成像模态下进行分析。结果发现,在图像解读任务中,Gemma、LLAVA、Qwen 和 MiniCPM 等模型在零样本下准确率仅为63%至75%,而少样本提示可将性能提升至86.8%,但仍存在误检和漏检现象,表明幻觉效应未完全消除。在文本到图像生成任务中,如 GPT-5 和 Gemini 2.5 Flash 等先进模型在宽松提示约束下,高达91%的生成场景存在生物学不合理性,揭示当前多模态大模型的根本性弱点。本研究系统评估了视觉推理与生成能力,为提升基于大模型的农业成像平台可靠性与可信度提供了关键洞见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are being rapidly adopted in agricultural imaging applications, ranging from crop interpretation to synthetic field image generation. However, these models frequently exhibit hallucinations outputs that appear confident yet deviate from biological or environmental reality potentially leading to misinformed agronomic insights. This study investigates such hallucinations in two complementary directions: image-to-text, where LLMs interpret crop or field imagery to describe conditions such as biotic and abiotic stresses, and text-to-image, where models generate synthetic agricultural scenes based on descriptive prompts. We examine errors involving biological inconsistency, contextual inaccuracy, and agronomic implausibility, evaluating the outputs under domain-informed criteria across multiple imaging modalities. Our analysis identifies recurring hallucination patterns within both interpretive and generative tasks. In image interpretation, LLMs (e.g., Gemma, LLAVA, Qwen, and MiniCPM) achieved modest zero-shot accuracy (63 to 75 percent), whereas few-shot prompting improved performance up to 86.8 percent, exhibiting false detections and missed infections, indicating residual hallucination effects. In text-to-image tasks, advanced models such as GPT-5 and Gemini 2.5 Flash generate up to 91 percent biologically inconsistent scenes under relaxed prompt constraints, revealing fundamental weaknesses in current LLMs. This systematic assessment of visual reasoning and generation offers critical insights toward enhancing the reliability and trustworthiness of LLM-based agricultural imaging platforms.

多模态幻觉农业图像大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。