用渗透测试发现音频深度伪造检测模型的漏洞
DeePen: Penetration Testing for Audio Deepfake Detection
- 不依赖目标模型,通过信号处理攻击评估检测系统脆弱性
- 时间拉伸和添加回声等简单操作可欺骗所有测试模型
- 部分攻击可通过针对性重训练缓解,但仍有持续有效的威胁
深度伪造——被篡改或伪造的音视频媒体——对个人、组织和社会构成重大安全风险。为应对这些挑战,通常采用基于机器学习的分类器来检测深度伪造内容。本文提出一种系统化的渗透测试方法DeePen,用于评估此类分类器的鲁棒性。该方法无需事先了解或访问目标深度伪造检测模型,而是通过一系列精心选择的信号处理修改(称为攻击)来探测模型漏洞。利用DeePen,我们分析了真实生产系统和公开学术模型检查点,结果表明所有测试系统均存在弱点,且可被时间拉伸或添加回声等简单操作可靠欺骗。此外,我们的研究发现,尽管某些攻击可通过知晓具体攻击方式后重新训练检测系统加以缓解,但其他攻击仍持续有效。
原文摘要 · Abstract (English)
Deepfakes - manipulated or forged audio and video media - pose significant security risks to individuals, organizations, and society at large. To address these challenges, machine learning-based classifiers are commonly employed to detect deepfake content. In this paper, we assess the robustness of such classifiers through a systematic penetration testing methodology, which we introduce as DeePen. Our approach operates without prior knowledge of or access to the target deepfake detection models. Instead, it leverages a set of carefully selected signal processing modifications - referred to as attacks - to evaluate model vulnerabilities. Using DeePen, we analyze both real-world production systems and publicly available academic model checkpoints, demonstrating that all tested systems exhibit weaknesses and can be reliably deceived by simple manipulations such as time-stretching or echo addition. Furthermore, our findings reveal that while some attacks can be mitigated by retraining detection systems with knowledge of the specific attack, others remain persistently effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。