arXiv:2601.05986cs.CVcs.CR2026-01

测试深度伪造检测模型在真实攻击下的鲁棒性,发现对抗训练效果因场景而异。

Deepfake detectors are DUMB: A benchmark to assess adversarial training robustness under transferability constraints

  • 构建跨数据集、跨模型的对抗攻击测试框架,模拟真实攻击环境。
  • 对抗训练在同分布下提升鲁棒性,但跨数据集时可能反而降低性能。
  • 提醒实际应用需根据场景选择防御策略,不能一概而论。

部署在真实环境中的深度伪造检测系统面临能够制造难以察觉扰动的攻击者,导致模型性能下降。尽管对抗训练是广泛采用的防御手段,但在攻击者知识有限且数据分布不匹配等现实条件下,其有效性仍缺乏深入研究。本文扩展了DUMB(数据源、模型架构与平衡)及DUMBer方法论,用于评估深度伪造检测器在转移性约束和跨数据集配置下的对抗攻击鲁棒性。实验涵盖五种前沿检测器(RECCE、SRM、XCeption、UCF、SPSL)、三种攻击方法(PGD、FGSM、FPBA)以及两个数据集(FaceForensics++ 和 Celeb-DF-V2)。从攻击者与防御者双重视角分析结果,并映射至分布不匹配场景。结果显示,对抗训练在同分布情况下可增强鲁棒性,但在跨数据集配置下,其效果取决于具体策略,可能反而削弱性能。研究强调在真实应用中需采用情境感知的防御策略。

原文摘要 · Abstract (English)

Deepfake detection systems deployed in real-world environments are subject to adversaries capable of crafting imperceptible perturbations that degrade model performance. While adversarial training is a widely adopted defense, its effectiveness under realistic conditions -- where attackers operate with limited knowledge and mismatched data distributions - remains underexplored. In this work, we extend the DUMB -- Dataset soUrces, Model architecture and Balance - and DUMBer methodology to deepfake detection. We evaluate detectors robustness against adversarial attacks under transferability constraints and cross-dataset configuration to extract real-world insights. Our study spans five state-of-the-art detectors (RECCE, SRM, XCeption, UCF, SPSL), three attacks (PGD, FGSM, FPBA), and two datasets (FaceForensics++ and Celeb-DF-V2). We analyze both attacker and defender perspectives mapping results to mismatch scenarios. Experiments show that adversarial training strategies reinforce robustness in the in-distribution cases but can also degrade it under cross-dataset configuration depending on the strategy adopted. These findings highlight the need for case-aware defense strategies in real-world applications exposed to adversarial attacks.

深度伪造对抗攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。