评测并缓解视频大模型对事件关系的幻觉问题
VERHallu: Evaluating and Mitigating Event Relation Hallucination in Video Large Language Models
- 构建新基准VERHallu,评估因果、时序、子事件关系
- 模型依赖先验知识,忽视帧级线索导致关系理解不全
- 提出关键帧传播策略,提升多事件推理且不降速
视频大语言模型存在多种幻觉。现有研究多关注事件、物体和场景的错认,却忽视事件关系幻觉。本文提出新型基准VERHallu,聚焦事件间的因果、时序与子事件关系,涵盖关系分类、问答及反事实问答三类任务,全面评估关系幻觉。该基准包含违背常见预训练分布的反直觉视频场景,每样本配有由人工标注的视觉-语言与纯语言偏差候选。分析显示,当前顶尖视频大模型在密集事件关系推理上表现不佳,常依赖先验知识而忽略帧级线索;虽能准确识别关键事件,却常遗漏周边子事件,导致关系理解不完整。为此,我们提出关键帧传播(KFP)策略,在中间层重分配帧级注意力,增强多事件理解。实验表明,该方法有效缓解事件关系幻觉,且不影响推理速度。
原文摘要 · Abstract (English)
Video Large Language Models (VideoLLMs) exhibit various types of hallucinations. Existing research has primarily focused on hallucinations involving the presence of events, objects, and scenes in videos, while largely neglecting event relation hallucination. In this paper, we introduce a novel benchmark for evaluating the Video Event Relation Hallucination, named VERHallu. This benchmark focuses on causal, temporal, and subevent relations between events, encompassing three types of tasks: relation classification, question answering, and counterfactual question answering, for a comprehensive evaluation of event relation hallucination. Additionally, it features counterintuitive video scenarios that deviate from typical pretraining distributions, with each sample accompanied by human-annotated candidates covering both vision-language and pure language biases. Our analysis reveals that current state-of-the-art VideoLLMs struggle with dense-event relation reasoning, often relying on prior knowledge due to insufficient use of frame-level cues. Although these models demonstrate strong grounding capabilities for key events, they often overlook the surrounding subevents, leading to an incomplete and inaccurate understanding of event relations. To tackle this, we propose a Key-Frame Propagating (KFP) strategy, which reallocates frame-level attention within intermediate layers to enhance multi-event understanding. Experiments show it effectively mitigates the event relation hallucination without affecting inference speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。