提出SHIELD框架,提升音频伪造检测在对抗攻击下的鲁棒性
SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks

- 通过协同学习融合输入输出,引入防御生成模型捕捉伪造痕迹
- 在三个数据集上将检测准确率从不足60%恢复至98%以上
- 适合安全敏感场景的语音验证系统部署使用
音频在语音识别、智能设备和会议系统中至关重要,但深度伪造等音频篡改技术带来严重信息误导风险。实证分析表明,现有深伪音频检测方法易受生成式反取证(AF)攻击影响,尤其针对基于生成对抗网络的攻击。本文提出新型协同学习方法SHIELD,以应对生成式AF攻击。通过引入辅助防御生成模型,实现输入与输出的协同学习,并设计三元组模型,利用辅助生成模型捕捉真实音频与被攻击音频之间的关联关系。所提SHIELD机制显著增强对生成式反取证攻击的防御能力,在多种生成模型下保持稳定性能。在三个数据集上,原方法检测准确率分别从95.49%降至59.77%(ASVspoof2019)、从99.44%降至38.45%(In-the-Wild)、从98.41%降至51.18%(HalfTruth)。而SHIELD在匹配与非匹配设置下,分别达到98.13%/98.78%(ASVspoof2019)、98.58%/98.62%(In-the-Wild)、99.57%/98.85%(HalfTruth)的平均准确率。
原文摘要 · Abstract (English)
Audio plays a crucial role in applications like speaker verification, voice-enabled smart devices, and audio conferencing. However, audio manipulations, such as deepfakes, pose significant risks by enabling the spread of misinformation. Our empirical analysis reveals that existing methods for detecting deepfake audio are often vulnerable to anti-forensic (AF) attacks, particularly those attacked using generative adversarial networks. In this article, we propose a novel collaborative learning method called SHIELD to defend against generative AF attacks. To expose AF signatures, we integrate an auxiliary generative model, called the defense (DF) generative model, which facilitates collaborative learning by combining input and output. Furthermore, we design a triplet model to capture correlations for real and AF attacked audios with real-generated and attacked-generated audios using auxiliary generative models. The proposed SHIELD strengthens the defense against generative AF attacks and achieves robust performance across various generative models. The proposed AF significantly reduces the average detection accuracy from 95.49% to 59.77% for ASVspoof2019, from 99.44% to 38.45% for In-the-Wild, and from 98.41% to 51.18% for HalfTruth for three different generative models. The proposed SHIELD mechanism is robust against AF attacks and achieves an average accuracy of 98.13%, 98.58%, and 99.57% in match, and 98.78%, 98.62%, and 98.85% in mismatch settings for the ASVspoof2019, In-the-Wild, and HalfTruth datasets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。