通过乘法攻击彻底清除生成模型痕迹,实现不可追溯的深度伪造。
Untraceable DeepFakes via Traceable Fingerprint Elimination
- 提出乘法攻击,从根源上消除生成模型留下的可追踪痕迹。
- 在9种生成模型上对6个先进溯源模型攻击成功率达97.08%,防御下仍超72.39%。
- 无需目标模型信息,通用性强,适合研究深度伪造防御漏洞者。
近期深度伪造溯源技术进步显著,能够提取生成模型(GMs)在图像中留下的痕迹,使深度伪造可被追溯至源模型。然而,现有规避攻击未能真正消除痕迹,仍易被防御手段抵消。本文发现,通过乘法攻击可从根本上消除生成模型痕迹,从而绕过即使增强的溯源模型。我们设计了一种仅使用真实数据训练的通用黑盒攻击方法,适用于多种生成模型且与溯源模型无关。实验表明,该方法在9种生成模型生成的深度伪造上,对6个先进溯源模型平均攻击成功率(ASR)达97.08%;即便存在防御机制,攻击成功率仍超过72.39%。本工作揭示了乘法攻击带来的潜在威胁,凸显了构建更鲁棒溯源模型的必要性。
原文摘要 · Abstract (English)
Recent advancements in DeepFakes attribution technologies have significantly enhanced forensic capabilities, enabling the extraction of traces left by generative models (GMs) in images, making DeepFakes traceable back to their source GMs. Meanwhile, several attacks have attempted to evade attribution models (AMs) for exploring their limitations, calling for more robust AMs. However, existing attacks fail to eliminate GMs' traces, thus can be mitigated by defensive measures. In this paper, we identify that untraceable DeepFakes can be achieved through a multiplicative attack, which can fundamentally eliminate GMs' traces, thereby evading AMs even enhanced with defensive measures. We design a universal and black-box attack method that trains an adversarial model solely using real data, applicable for various GMs and agnostic to AMs. Experimental results demonstrate the outstanding attack capability and universal applicability of our method, achieving an average attack success rate (ASR) of 97.08\% against 6 advanced AMs on DeepFakes generated by 9 GMs. Even in the presence of defensive mechanisms, our method maintains an ASR exceeding 72.39\%. Our work underscores the potential challenges posed by multiplicative attacks and highlights the need for more robust AMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。