首个针对多图幻觉的评测与缓解方案,提升模型跨图理解准确性
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
- 构建多图幻觉评测基准MIHBench,涵盖存在、数量、身份三类任务
- 发现图像数量越多、单图幻觉越强,且负样本位置影响身份一致性判断
- 提出动态注意力平衡机制,显著降低多图场景下幻觉率
尽管多模态大语言模型中的幻觉问题受到广泛关注,但现有研究主要集中在单图场景,多图情况下的幻觉仍缺乏系统研究。为此,我们首次系统性地研究多图场景下幻觉现象,并提出MIHBench——一个专门用于评估跨多图对象相关幻觉的基准。该基准包含三个核心任务:多图对象存在幻觉、多图对象计数幻觉以及对象身份一致性幻觉,分别针对对象存在性、数量推理和跨视角身份一致性等语义理解能力。通过大规模实验,我们发现:图像输入数量与幻觉发生概率呈正相关;单图幻觉倾向与多图幻觉高度相关;相同对象图像比例及负样本在图像序列中的位置显著影响身份一致性幻觉。为应对上述挑战,我们提出动态注意力平衡机制,在保持整体视觉注意力比例的同时调整跨图注意力分布。在多个先进多模态大模型上的实验表明,该方法有效减少幻觉发生,提升多图场景下的语义整合与推理稳定性。
原文摘要 · Abstract (English)
Despite growing interest in hallucination in Multimodal Large Language Models, existing studies primarily focus on single-image settings, leaving hallucination in multi-image scenarios largely unexplored. To address this gap, we conduct the first systematic study of hallucinations in multi-image MLLMs and propose MIHBench, a benchmark specifically tailored for evaluating object-related hallucinations across multiple images. MIHBench comprises three core tasks: Multi-Image Object Existence Hallucination, Multi-Image Object Count Hallucination, and Object Identity Consistency Hallucination, targeting semantic understanding across object existence, quantity reasoning, and cross-view identity consistency. Through extensive evaluation, we identify key factors associated with the occurrence of multi-image hallucinations, including: a progressive relationship between the number of image inputs and the likelihood of hallucination occurrences; a strong correlation between single-image hallucination tendencies and those observed in multi-image contexts; and the influence of same-object image ratios and the positional placement of negative samples within image sequences on the occurrence of object identity consistency hallucination. To address these challenges, we propose a Dynamic Attention Balancing mechanism that adjusts inter-image attention distributions while preserving the overall visual attention proportion. Experiments across multiple state-of-the-art MLLMs demonstrate that our method effectively reduces hallucination occurrences and enhances semantic integration and reasoning stability in multi-image scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。