arXiv:2511.20295cs.CV2025-11中稿 · CVPR被引 1

提出视频反事实解释框架,让模型决策过程更透明。

Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations

  • 通过优化初始噪声和分阶段搜索,生成时序连贯的视频反事实样本。
  • 在多个数据集上生成的反事实视频与原视频相似且真实可信。
  • 适合关注视频模型可解释性的研究者与工程师使用。

反事实解释(CFE)是模型输入的最小且语义有意义的修改,能改变模型预测结果,揭示模型依赖的关键特征,提供对比性解释。当前最先进的视觉反事实方法主要针对图像分类器,对视频模型的解释仍缺乏探索。有效的视频反事实解释需具备物理合理性、时间连贯性和平滑运动轨迹。现有基于图像的反事实方法无法生成时序一致、平滑且物理合理的视频反事实。为此,我们提出 Back To The Feature (BTTF) 优化框架,引入两项新特性:1)基于输入视频首帧条件化的初始潜在噪声检索机制;2)两阶段优化策略,用于在输入视频邻域内搜索反事实视频。两项优化均仅由目标分类器引导,确保解释忠实性。为加速收敛,还引入渐进式优化策略,逐步增加去噪步数。在 Shape-Moving(运动分类)、MEAD(情绪分类)和 NTU RGB+D(动作分类)等视频数据集上的大量实验表明,BTTF 能有效生成视觉相似、真实且有效的反事实视频,为理解分类器决策机制提供具体洞察。

原文摘要 · Abstract (English)

Counterfactual explanations (CFEs) are minimal and semantically meaningful modifications of the input of a model that alter the model predictions. They highlight the decisive features the model relies on, providing contrastive interpretations for classifiers. State-of-the-art visual counterfactual explanation methods have primarily focused on interpreting image classifiers, leaving the domain of video models relatively underexplored. For the video CFEs to be useful, they have to be physically plausible, temporally coherent, and exhibit smooth motion trajectories. Existing CFE image-based methods, designed to explain image classifiers, lack the capacity to generate temporally coherent, smooth and physically plausible video CFEs. To address this, we propose Back To The Feature (BTTF), an optimization framework that generates video CFEs. Our method introduces two novel features, 1) an optimization scheme to retrieve the initial latent noise conditioned by the first frame of the input video, 2) a two-stage optimization strategy to enable the search for counterfactual videos in the vicinity of the input video. Both optimization processes are guided solely by the target classifier, ensuring the explanation is faithful. To accelerate convergence, we also introduce a progressive optimization strategy that incrementally increases the number of denoising steps. Extensive experiments on video datasets such as Shape-Moving (motion classification), MEAD (emotion classification), and NTU RGB+D (action classification) show that our BTTF effectively generates valid, visually similar and realistic counterfactual videos that provide concrete insights into the classifier's decision-making mechanism.

视频解释反事实可解释性扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。