arXiv:2506.07184cs.AIcs.CL2025-06被引 3

提出新方法减轻多模态模型在序列图像中的行为幻觉。

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images

  • 通过自适应时序窗口检测幻觉,再用正交投影修正。
  • 在标准数据集上使行为幻觉减少超10%。
  • 适合需要可靠图像序列理解的场景。

多模态大模型虽在多项任务中表现优异,但仍存在幻觉问题,限制其可靠性与广泛应用。现有研究多关注客观幻觉,但针对序列图像,还存在行为幻觉,该问题较少被研究。本文揭示行为幻觉主要源于先验偏差与雪球效应。为此,提出SHE(Sequence Hallucination Eradication)轻量级两阶段框架:(1) 利用自适应时序窗口进行视觉-文本对齐检测;(2) 通过正交投影至联合嵌入空间缓解幻觉。同时提出新度量指标BEACH以量化行为幻觉严重程度。在标准基准上的实验证明,SHE在保持描述准确性的同时,使行为幻觉在BEACH指标上降低超过10%。

原文摘要 · Abstract (English)

While multimodal large language models excel at various tasks, they still suffer from hallucinations, which limit their reliability and scalability for broader domain applications. To address this issue, recent research mainly focuses on objective hallucination. However, for sequential images, besides objective hallucination, there is also behavioral hallucination, which is less studied. This work aims to fill in the gap. We first reveal that behavioral hallucinations mainly arise from two key factors: prior-driven bias and the snowball effect. Based on these observations, we introduce SHE (Sequence Hallucination Eradication), a lightweight, two-stage framework that (1) detects hallucinations via visual-textual alignment check using our proposed adaptive temporal window and (2) mitigates them via orthogonal projection onto the joint embedding space. We also propose a new metric (BEACH) to quantify behavioral hallucination severity. Empirical results on standard benchmarks demonstrate that SHE reduces behavioral hallucination by over 10% on BEACH while maintaining descriptive accuracy.

多模态幻觉序列图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。