用记忆提升视频事件理解,但需防虚假记忆干扰
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
- 利用历史视频事件作为上下文记忆增强理解
- 虚假记忆导致错误推理,性能下降约15%
- 提出抗幻觉记忆修正方法,适合视频分析研究者
多模态大语言模型在整体视频理解上表现优异,但在流式视频处理(将视频视为视觉事件序列)方面仍待深入。直觉上,利用过往事件作为记忆可增强当前事件的上下文与时间理解。本文发现,借助记忆确实能提升事件理解能力,但因记忆依赖前序事件的预测,可能包含错误信息,引发幻觉并导致性能下降。为此,我们提出一种抗幻觉的记忆修正方法,有效缓解虚假记忆对事件理解的影响。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have demonstrated strong performance in understanding videos holistically, yet their ability to process streaming videos-videos are treated as a sequence of visual events-remains underexplored. Intuitively, leveraging past events as memory can enrich contextual and temporal understanding of the current event. In this paper, we show that leveraging memories as contexts helps MLLMs better understand video events. However, because such memories rely on predictions of preceding events, they may contain misinformation, leading to confabulation and degraded performance. To address this, we propose a confabulation-aware memory modification method that mitigates confabulated memory for memory-enhanced event understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。