arXiv:2409.17647cs.CV2024-09NeurIPS被引 34

提出多事件因果发现任务,解析长视频中多个事件的因果链。

MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning

  • 基于格兰杰因果思想,用掩码预测模型测试事件间的因果关系。
  • 在长视频多事件场景下,性能优于GPT-4o和VideoLLaVA 5.7%与4.1%。
  • 适用于需要深度理解事件演变逻辑的视频分析任务。

视频因果推理旨在从因果视角实现对视频内容的高层理解。然而,现有视频推理任务范围有限,主要以问答形式进行,集中在仅含单一事件和简单因果关系的短视频上,缺乏对多事件长视频的全面、结构化因果分析。为填补这一空白,我们提出新任务与数据集——多事件因果发现(MECD),旨在揭示长视频中跨时间分布的多个事件之间的因果关系。给定事件的视觉片段和文本描述,MECD要求识别事件间的因果关联,构建解释最终结果事件发生原因与机制的结构化事件级因果图。为此,我们设计了一种受格兰杰因果启发的新框架,采用高效的掩码事件预测模型执行事件格兰杰检验,通过对比掩码前提事件与未掩码时对结果事件的预测差异来估计因果性。此外,引入前门调整与反事实推理等因果推断技术,应对因果混淆与虚假因果等挑战。实验表明,该框架在多事件视频因果推理中有效,相较GPT-4o和VideoLLaVA分别提升5.7%和4.1%。

原文摘要 · Abstract (English)

Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily executed in a question-answering paradigm and focusing on short videos containing only a single event and simple causal relationships, lacking comprehensive and structured causality analysis for videos with multiple events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relationships between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD requires identifying the causal associations between these events to derive a comprehensive, structured event-level video causal diagram explaining why and how the final result event occurred. To address MECD, we devise a novel framework inspired by the Granger Causality method, using an efficient mask-based event prediction model to perform an Event Granger Test, which estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to address challenges in MECD like causality confounding and illusory causality. Experiments validate the effectiveness of our framework in providing causal relationships in multi-event videos, outperforming GPT-4o and VideoLLaVA by 5.7% and 4.1%, respectively.

因果发现视频理解长视频多事件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。