黑客可植入隐蔽后门,让深伪检测器失效
Where the Devil Hides: Deepfake Detectors Can No Longer Be Trusted
- 设计隐形触发器,控制深伪检测器行为
- 在污染数据中注入后门,使检测器误判特定样本
- 攻击者可绕过检测,适合安全与防御研究者
随着AI生成技术发展,深伪人脸已高度逼真,难以被肉眼识别。为此,基于深度神经网络的深伪检测器应运而生,通常依赖第三方数据集进行训练。然而,这一流程存在严重安全隐患:若第三方数据提供方恶意注入被污染的数据,训练出的检测器将被植入“后门”,在遇到含特定触发器的样本时产生异常行为。此类触发器可能被出售给攻击者,使其规避检测、逃避追责。本文深入研究该风险,提出一种可生成可控、语义抑制、自适应且不可见触发模式的触发器生成器,并探讨脏标签与干净标签两种污染场景实现后门注入。大量实验表明,该方法在有效性、隐蔽性和实用性上均优于多个基线。
原文摘要 · Abstract (English)
With the advancement of AI generative techniques, Deepfake faces have become incredibly realistic and nearly indistinguishable to the human eye. To counter this, Deepfake detectors have been developed as reliable tools for assessing face authenticity. These detectors are typically developed on Deep Neural Networks (DNNs) and trained using third-party datasets. However, this protocol raises a new security risk that can seriously undermine the trustfulness of Deepfake detectors: Once the third-party data providers insert poisoned (corrupted) data maliciously, Deepfake detectors trained on these datasets will be injected ``backdoors'' that cause abnormal behavior when presented with samples containing specific triggers. This is a practical concern, as third-party providers may distribute or sell these triggers to malicious users, allowing them to manipulate detector performance and escape accountability. This paper investigates this risk in depth and describes a solution to stealthily infect Deepfake detectors. Specifically, we develop a trigger generator, that can synthesize passcode-controlled, semantic-suppression, adaptive, and invisible trigger patterns, ensuring both the stealthiness and effectiveness of these triggers. Then we discuss two poisoning scenarios, dirty-label poisoning and clean-label poisoning, to accomplish the injection of backdoors. Extensive experiments demonstrate the effectiveness, stealthiness, and practicality of our method compared to several baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。