测试发现视频大模型会把广告片段错当成主视频内容,导致错误关联。
Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models

- 用插入广告片段的方法测试模型对时间连贯性的理解
- 11个主流视频模型均出现主体与事件错配的系统性幻觉
- 适合关注视频理解可靠性与模型缺陷的研究者阅读
视频理解的关键能力是跨时间可靠关联主体与事件,但视频大语言模型(VideoLLMs)是否真正具备此能力尚不明确。本文提出DistractionBench评测框架,通过在长视频中插入短广告片段等可控干扰,检验模型在干扰下的主体-事件关联能力。实验发现,所有11个主流VideoLLMs均频繁将广告片段中的动作错误归因于主视频主体,产生虚假交互。这种系统性错误被定义为‘袋中事件’(bag-of-events, BoE)行为,即模型将视频视为事件集合而非时序序列。结果表明,当前VideoLLMs缺乏可靠的时序定位机制,亟需发展更稳健的主体-事件关联模型。
原文摘要 · Abstract (English)
A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In this work, we introduce DistractionBench to evaluate whether VideoLLMs can robustly link subjects and events in the presence of unrelated video segments. Through controlled interventions, such as inserting short advertisement clips into longer videos, we show that VideoLLMs frequently hallucinate interactions between entities from different segments, incorrectly attributing actions from injected advertisements to subjects in the main video. We characterize this systematic hallucination as bag-of-events (BoE) behavior, where models process videos as collections of events rather than temporally structured sequences. Evaluating 11 popular VideoLLMs, we find that all models exhibit substantial BoE behavior. Our findings suggest that VideoLLMs lack reliable mechanisms for temporal grounding and motivate the development of models with more robust subject-event association.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。