首个针对长视频语义聚合幻觉的基准,揭示模型在复杂事件中误判的深层原因。
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- 构建首个长视频幻觉评测基准,聚焦多帧语义整合过程中的错误
- 发现幻觉率随语义复杂度和快速变化而上升,达27.7%显著降低
- 提出位置编码与强化学习策略,提升跨事件语义区分能力
视频多模态大模型在视频理解方面取得显著进展,但仍易产生与视频内容不符或无关的幻觉。现有幻觉评测多集中于短视频,归因于语言先验、缺失帧或视觉-语言偏差。然而这些因素无法解释长视频中出现的特定幻觉:模型输出看似合理,但帧级语义被错误聚合为事件级错误结论。我们称之为语义聚合幻觉(SAH),其源于多帧语义向事件级语义的整合过程。由于长视频涉及更多复杂事件,此类幻觉尤为突出。为此,我们提出首个专注于长视频的基准ELV-Halluc,系统研究SAH。实验验证了其存在性,并发现幻觉率随语义复杂度上升;模型在语义快速变化时更易犯错。我们提出采用位置编码缓解该问题,并结合直接偏好优化(DPO)增强模型对事件内与事件间语义的区分能力。构建包含8,000组对抗样本的数据集,在ELV-Halluc与Video-MME上均实现性能提升,其中SAH比例降低27.7%。
原文摘要 · Abstract (English)
Video multimodal large language models (Video-MLLMs) have achieved remarkable progress in video understanding. However, they remain vulnerable to hallucination-producing content inconsistent with or unrelated to video inputs. Previous video hallucination benchmarks primarily focus on short-videos. They attribute hallucinations to factors such as strong language priors, missing frames, or vision-language biases introduced by the visual encoder. While these causes indeed account for most hallucinations in short videos, they still oversimplify the cause of hallucinations. Sometimes, models generate incorrect outputs but with correct frame-level semantics. We refer to this type of hallucination as Semantic Aggregation Hallucination (SAH), which arises during the process of aggregating frame-level semantics into event-level semantic groups. Given that SAH becomes particularly critical in long videos due to increased semantic complexity across multiple events, it is essential to separate and thoroughly investigate the causes of this type of hallucination. To address the above issues, we introduce ELV-Halluc, the first benchmark dedicated to long-video hallucination, enabling a systematic investigation of SAH. Our experiments confirm the existence of SAH and show that it increases with semantic complexity. Additionally, we find that models are more prone to SAH on rapidly changing semantics. Moreover, we discuss potential approaches to mitigate SAH. We demonstrate that positional encoding strategy contributes to alleviating SAH, and further adopt DPO strategy to enhance the model's ability to distinguish semantics within and across events. To support this, we curate a dataset of 8K adversarial data pairs and achieve improvements on both ELV-Halluc and Video-MME, including a substantial 27.7% reduction in SAH ratio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。