针对多模态模型幻觉问题,设计可定位原因的系统化测试基准。
ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation

- 通过对抗图像和诱导性提问,设计四类针对性任务。
- 发现当前多模态模型对特定诱因仍高度敏感,准确率普遍低于60%。
- 适合研究模型鲁棒性、幻觉成因及评测方法的学者使用。
尽管多模态大语言模型(MLLMs)在视觉-语言理解方面取得了快速进展,但仍易产生与视觉输入不一致的幻觉。现有评测基准多聚焦于幻觉结果的检测,而非其根本成因。此外,许多基准依赖简单场景和有限评估形式,已无法有效挑战前沿模型。为此,我们提出 ReactBench,一个以原因驱动的幻觉评测基准,包含多任务与考试式评估格式。通过生成对抗图像和诱发幻觉的查询,该基准设计了四类任务:关系消解、反事实属性、变化溯源与密集计数,系统暴露共现偏差、语言先验、跨图比较感知缺陷及细粒度感知瓶颈。除标准准确率外,还采用思维链(Chain-of-Thought)分析各任务中幻觉的细粒度子原因。大量评估显示,当前 MLLMs 对特定诱因仍显著脆弱,验证了 ReactBench 在诊断与提升多模态模型鲁棒性方面的系统性与可解释性价值。项目页面见 https://reactbench.github.io/。
原文摘要 · Abstract (English)
While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinations, producing responses that are inconsistent with the visual input. Existing benchmarks predominantly focus on detecting hallucination outcomes rather than evaluating the underlying causes of these failures. Moreover, many benchmarks rely on simplistic scenarios and limited evaluation formats that no longer challenge state-of-the-art models. To address these limitations, we introduce ReactBench, a cause-driven hallucination benchmark featuring multiple tasks and an exam-style evaluation format. By generating adversarial images and hallucination-inducing queries, ReactBench introduces four targeted tasks: Relational Erasure, Counterfactual Attribute, Alteration Tracing, and Dense Counting. These tasks systematically expose co-occurrence bias, language priors, cross-image comparative perception deficiencies, and fine-grained perceptual bottlenecks. Beyond standard accuracy-based evaluation, we leverage Chain-of-Thought reasoning to identify fine-grained sub-causes of hallucination within each task. Extensive evaluations reveal that current MLLMs remain notably vulnerable to cause-specific hallucination triggers, demonstrating the value of ReactBench as a systematic and interpretable testbed for diagnosing and improving multimodal model robustness. The project page is available at https://reactbench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。