arXiv:2606.09181cs.CVcs.LG2026-06

提出因果推理框架,让视频问答更准确地识别关键证据。

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA

论文配图:Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA
图 1 · 摘自论文原文
  • 用因果模型分离视觉线索与干扰因素
  • 在多个数据集上提升答案准确率和可靠性
  • 适合需要可解释视频理解的场景

近年来,视频多模态模型显著提升了VideoQA性能,但这些系统常依赖虚假统计关联而非真实因果证据,导致推理不忠实且脆弱,尤其在复杂现实场景中表现不佳。现有方法或依赖跨模态相关性、需昂贵标注资源,或因果假设不足,且通常仅在时间区间层面操作,难以显式分离因果视觉线索与混杂因素,证据定位粒度有限。为此,我们提出基于反事实推理的细粒度证据解耦框架CREDiT。CREDiT采用结构化因果模型建模VideoQA过程,在独立性和最小性约束下,学习显式分解为因果与非因果成分的跨模态表示。为实现可信解耦,引入特征级因果干预,构建近似因果效应的反事实输入,同时抑制非因果相关性。在NExT-GQA、SportsQA和SPORTU-video上的大量实验表明,CREDiT在通用与复杂体育场景中均持续提升答案准确率与推理可靠性,使VideoQA系统更具可信度。

原文摘要 · Abstract (English)

Recent advances in video multimodal models have significantly improved VideoQA performance. However, these systems often rely on spurious statistical correlations rather than answer-relevant causal evidence, resulting in unfaithful and brittle reasoning, especially in complex real-world scenarios. Existing methods either rely on cross-modality correlations, costly curated training resources, or insufficient causal assumptions and constraints, and typically operate at the time-interval level. As a result, they fail to explicitly disentangle causal visual cues from confounders and provide limited fine-grained evidence localization. To address this issue, we propose a Counterfactual Reasoning framework for fine-grained Evidence Disentanglement (CREDiT). CREDiT formulates the VideoQA process using a structural causal model and learns cross-modality representations that are explicitly decomposed into causal and non-causal components under independence and minimality constraints. To facilitate faithful disentanglement, we introduce feature-level causal interventions and construct counterfactual inputs that approximate causal effects while suppressing non-causal correlations. Extensive experiments on NExT-GQA, SportsQA, and SPORTU-video demonstrate that CREDiT consistently improves answer accuracy and reasoning reliability across both generic and complex sports scenarios, leading to more trustworthy VideoQA systems.

视频问答因果推理证据解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。