arXiv:2502.03957cs.CVcs.AI2025-02中稿 · publication, AI4MF…被引 8

用对抗样本生成更精准的伪造图像解释,提升检测器可解释性。

Improving the Perturbation-Based Explanation of Deepfake Detectors Through the Use of Adversarially-Generated Samples

  • 基于自然进化策略生成对抗样本,反向扰动以还原真实图像。
  • 在FaceForensics++数据集上,改进后解释方法定位篡改区域更准确。
  • 适用于需要高精度视觉解释的深度伪造检测场景。

本文提出利用对抗生成的输入图像样本来构建扰动掩码,用于推断不同输入特征的重要性并生成可视化解释。这些样本通过自然进化策略生成,旨在使原本被判定为伪造的图像被检测器分类为真实。该方法应用于四种基于扰动的解释方法(LIME、SHAP、SOBOL 和 RISE),并在 SOTA 深度伪造检测模型、基准数据集 FaceForensics++ 及对应解释评估框架下进行评估。定量分析表明,所提扰动方法显著提升了解释方法性能;定性分析显示,改进后的解释方法能更精确地勾勒出图像中被篡改区域,从而提供更具实用价值的解释。

原文摘要 · Abstract (English)

In this paper, we introduce the idea of using adversarially-generated samples of the input images that were classified as deepfakes by a detector, to form perturbation masks for inferring the importance of different input features and produce visual explanations. We generate these samples based on Natural Evolution Strategies, aiming to flip the original deepfake detector's decision and classify these samples as real. We apply this idea to four perturbation-based explanation methods (LIME, SHAP, SOBOL and RISE) and evaluate the performance of the resulting modified methods using a SOTA deepfake detection model, a benchmarking dataset (FaceForensics++) and a corresponding explanation evaluation framework. Our quantitative assessments document the mostly positive contribution of the proposed perturbation approach in the performance of explanation methods. Our qualitative analysis shows the capacity of the modified explanation methods to demarcate the manipulated image regions more accurately, and thus to provide more useful explanations.

深度伪造可解释性对抗样本特征重要性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。