arXiv:2501.07227cs.CV2025-01TPAMI被引 8

提出视频事件级因果图发现新任务,实现长视频多事件因果推理。

MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning

  • 基于格兰杰因果思想设计掩码预测框架,通过对比事件遮蔽前后预测差异判断因果关系。
  • 在长视频中成功识别多事件因果链,相较GPT-4o提升5.77%,视频问答准确率显著提高。
  • 适合关注视频因果理解、多事件推理与可解释性模型的研究者和应用开发者。

视频因果推理旨在从因果视角实现对视频的高层理解,但现有方法受限于问题回答范式,仅关注短片段中的孤立事件与简单因果关系,缺乏对包含多个相互关联事件的长视频进行系统性因果分析的能力。为此,本文提出新任务与数据集——多事件因果发现(MECD),目标是揭示长视频中跨时间分布的事件之间的因果关系。给定视觉片段与事件文本描述,MECD旨在识别事件间的因果关联,构建完整且结构化的事件级视频因果图,以解释结果事件发生的原因与过程。为应对挑战,本文设计了一种受格兰杰因果启发的新框架,引入基于掩码的事件预测模型进行事件格兰杰检验,通过比较前提事件被遮蔽与未遮蔽时对结果事件的预测差异来估计因果性。同时结合前门调整与反事实推理等因果推断技术,缓解混杂与虚假因果问题。此外,引入上下文链推理增强泛化能力。实验表明,该框架在完整因果关系推理上优于GPT-4o和VideoChat2,分别提升5.77%和2.70%。进一步实验显示,所生成的因果图可有效辅助下游视频理解任务,如视频问答与事件预测。

原文摘要 · Abstract (English)

Video causal reasoning aims to achieve a high-level understanding of videos from a causal perspective. However, it exhibits limitations in its scope, primarily executed in a question-answering paradigm and focusing on brief video segments containing isolated events and basic causal relations, lacking comprehensive and structured causality analysis for videos with multiple interconnected events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relations between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD identifies the causal associations between these events to derive a comprehensive and structured event-level video causal graph explaining why and how the result event occurred. To address the challenges of MECD, we devise a novel framework inspired by the Granger Causality method, incorporating an efficient mask-based event prediction model to perform an Event Granger Test. It estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to mitigate challenges in MECD like causality confounding and illusory causality. Additionally, context chain reasoning is introduced to conduct more robust and generalized reasoning. Experiments validate the effectiveness of our framework in reasoning complete causal relations, outperforming GPT-4o and VideoChat2 by 5.77% and 2.70%, respectively. Further experiments demonstrate that causal relation graphs can also contribute to downstream video understanding tasks such as video question answering and video event prediction.

因果推理视频理解长视频图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。