通过因果反演分析视频事件,定位导致目标事件的根源触发因素。
Finding the Trigger: Causal Abductive Reasoning on Video Events
- 构建时空语义双通道关系网络,挖掘视频事件间因果关联。
- 在合成与真实视频数据集上验证,有效识别关键触发事件。
- 适用于安防监控、影视内容管理等需溯源分析的场景。
本文提出视频事件因果反演新任务(CARVE),旨在识别视频中事件间的因果关系,并生成解释目标事件发生的因果链假设。为推动该方向研究,我们构建了两个新基准数据集,包含合成与真实视频,其触发-目标标签通过创新的反事实合成方法生成。针对挑战,提出因果事件关系网络(CERN),在时序与语义空间中分析事件间关系,高效识别根源触发事件。大量实验表明,事件关系表征学习与交互建模在解决视频因果推理问题中至关重要。CARVE任务、配套数据集及CERN框架将显著推进视频因果推理研究,并助力视频监控、根因分析与影视内容管理等应用。
原文摘要 · Abstract (English)
This paper introduces a new problem, Causal Abductive Reasoning on Video Events (CARVE), which involves identifying causal relationships between events in a video and generating hypotheses about causal chains that account for the occurrence of a target event. To facilitate research in this direction, we create two new benchmark datasets with both synthetic and realistic videos, accompanied by trigger-target labels generated through a novel counterfactual synthesis approach. To explore the challenge of solving CARVE, we present a Causal Event Relation Network (CERN) that examines the relationships between video events in temporal and semantic spaces to efficiently determine the root-cause trigger events. Through extensive experiments, we demonstrate the critical roles of event relational representation learning and interaction modeling in solving video causal reasoning challenges. The introduction of the CARVE task, along with the accompanying datasets and the CERN framework, will advance future research on video causal reasoning and significantly facilitate various applications, including video surveillance, root-cause analysis and movie content management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。