用潜在扩散模型生成视频反事实解释,提升可解释性与实时性。
LD-ViCE: Latent Diffusion Model for Video Counterfactual Explanations

- 在潜在空间中使用扩散模型生成反事实视频,降低计算开销。
- 在3个数据集上优于现有方法,回归准确率显著提升且时序一致性高。
- 适合需要理解视频AI决策过程的研究者和开发者。
基于视频的AI系统在自动驾驶、医疗等关键领域应用日益广泛,但其决策解释仍因视频数据的时空复杂性和深度学习模型的黑箱特性而困难。现有解释方法常缺乏时序连贯性,且无法提供可操作的因果洞察。当前反事实解释方法通常未融合目标模型的指导,导致语义保真度不足。本文提出潜扩散视频反事实解释框架LD-ViCE,通过在潜在空间中使用先进扩散模型,显著降低解释生成的计算成本,并通过额外优化步骤生成真实且可解释的反事实视频。在三个不同视频数据集(EchoNet-Dynamic,FERV39k,Something-Something V2)上,针对分类与回归任务的多种目标模型进行实验,结果表明LD-ViCE具有良好泛化能力并达到当前最优性能。在EchoNet-Dynamic数据集上,其回归准确率显著高于已有方法,且表现出高时序一致性;优化阶段进一步提升感知质量。定性分析证实,LD-ViCE生成的解释具有语义意义且时序连贯,为模型行为提供可行动的洞察。该方法通过视觉一致的反事实解释,提升了视频类AI系统的可信度与可解释性。
原文摘要 · Abstract (English)
Video-based AI systems are increasingly adopted in safety-critical domains such as autonomous driving and healthcare. However, interpreting their decisions remains challenging due to the inherent spatiotemporal complexity of video data and the opacity of deep learning models. Existing explanation techniques often suffer from limited temporal coherence and a lack of actionable causal insights. Current counterfactual explanation methods typically do not incorporate guidance from the target model, reducing semantic fidelity and practical utility. We introduce Latent Diffusion for Video Counterfactual Explanations (LD-ViCE), a novel framework designed to explain the behavior of video-based AI models. Compared to previous approaches, LD-ViCE reduces the computational costs of generating explanations by operating in latent space using a state-of-the-art diffusion model, while producing realistic and interpretable counterfactuals through an additional refinement step. Experiments on three diverse video datasets - EchoNet-Dynamic (cardiac ultrasound), FERV39k (facial expression), and Something-Something V2 (action recognition) with multiple target models covering both classification and regression tasks, demonstrate that LD-ViCE generalizes well and achieves state-of-the-art performance. On the EchoNet-Dynamic dataset, LD-ViCE achieves significantly higher regression accuracy than prior methods and exhibits high temporal consistency, while the refinement stage further improves perceptual quality. Qualitative analyses confirm that LD-ViCE produces semantically meaningful and temporally coherent explanations, providing actionable insights into model behavior. LD-ViCE advances the trustworthiness and interpretability of video-based AI systems through visually coherent counterfactual explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。