arXiv:2504.15479cs.LGcs.CV2025-04被引 1

用潜在空间对抗攻击生成可解释的反事实图像,同时量化特征重要性。

Unifying Image Counterfactuals and Feature Attributions with Latent-Space Adversarial Attacks

  • 在低维流形上对图像表示做对抗攻击,生成反事实图像。
  • 结合辅助数据集,量化原始与反事实图像间的特征变化,生成全局解释。
  • 方法高效且兼容现代生成模型,适用于图像分类解释任务。

反事实解释是理解机器学习预测的重要框架,但对计算机视觉模型而言难以生成:传统基于梯度的方法易产生对抗样本,即像素的微小扰动导致预测大幅变化。本文提出一种新方法——反事实攻击,通过在图像表示的低维流形上进行对抗攻击,生成反事实图像。该方法易于实现,且能灵活适配当前生成建模技术。此外,利用额外的图像描述符数据集,可为反事实图像附加特征归因,量化原始与反事实图像之间的差异。这些重要性分数可聚合为全局反事实解释,揭示驱动模型预测的核心特征。该统一机制适用于任何反事实方法,但本方法具有显著计算效率优势。我们在MNIST和CelebA数据集上验证了其有效性。

原文摘要 · Abstract (English)

Counterfactuals are a popular framework for interpreting machine learning predictions. These what if explanations are notoriously challenging to create for computer vision models: standard gradient-based methods are prone to produce adversarial examples, in which imperceptible modifications to image pixels provoke large changes in predictions. We introduce a new, easy-to-implement framework for counterfactual images that can flexibly adapt to contemporary advances in generative modeling. Our method, Counterfactual Attacks, resembles an adversarial attack on the representation of the image along a low-dimensional manifold. In addition, given an auxiliary dataset of image descriptors, we show how to accompany counterfactuals with feature attribution that quantify the changes between the original and counterfactual images. These importance scores can be aggregated into global counterfactual explanations that highlight the overall features driving model predictions. While this unification is possible for any counterfactual method, it has particular computational efficiency for ours. We demonstrate the efficacy of our approach with the MNIST and CelebA datasets.

反事实解释特征归因对抗攻击图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。