测试10种音频伪造检测模型在18类真实噪声下的表现,发现压缩和修改最致命。
Measuring the Robustness of Audio Deepfake Detection under Real-World Corruption
- 系统测试10种模型对18类噪声、修改、压缩的抗干扰能力
- 语音基础模型普遍优于传统模型,大模型更鲁棒但收益递减
- 训练时加数据增强或推理时做语音增强可提升未知噪声应对力
深度伪造已成为生成式AI中广泛且迅速加剧的问题,涵盖图像、音频和视频。其中,音频深度伪造尤为令人担忧,因为高质量语音合成工具日益普及,且合成语音极易通过社交媒体和自动电话传播。因此,检测音频深度伪造对防止人工智能生成语音的滥用至关重要。然而,现实中的音频常受噪声、音频修改和压缩等污染,会显著降低检测性能。本文系统评估了10种音频深度伪造检测模型在18种常见污染类型下的鲁棒性,这些污染分为三类:噪声扰动、音频修改和压缩。使用传统深度学习模型和最新的语音基础模型,研究得出四个关键发现:(1)大多数模型对噪声具有鲁棒性,但对音频修改和压缩仍脆弱,尤其是神经编码器;(2)语音基础模型在多数污染场景下均优于传统模型,可能得益于在多样化音频数据集上的大规模预训练;(3)模型规模增大可提升鲁棒性,但收益随模型变大而递减;(4)通过训练时有针对性的数据增强或推理时进行语音增强,可改善对未见过污染的鲁棒性。这些发现强调了在多样真实污染条件下评估音频深度伪造检测器的重要性,并呼吁开发更具鲁棒性的检测框架以实现实际部署。我们进一步倡导,未来所有媒体的深度伪造检测研究都应考虑现实环境中多样且不可预测的失真。
原文摘要 · Abstract (English)
Deepfakes have emerged as a widespread and rapidly escalating concern in generative AI, spanning images, audio, and videos. Among these, audio deepfakes are particularly alarming due to the growing accessibility of high-quality voice synthesis tools and the ease with which synthetic speech can be distributed through social media and robocalls. Consequently, detecting audio deepfakes is critical for combating the misuse of AI-generated speech. However, real-world audio is often affected by corruptions such as noise, audio modification, and compression, which can significantly degrade detection performance. In this work, we systematically evaluate the robustness of 10 audio deepfake detection models against 18 common corruption types, grouped into three categories: noise perturbation, audio modification, and compression. Using both traditional deep learning models and state-of-the-art speech foundation models, our study yields four key insights. (1) Most models are robust to noise but remain vulnerable to audio modifications and compression, especially neural codecs. (2) Speech foundation models consistently outperform traditional models across most corruption scenarios, likely due to large-scale pre-training on diverse audio datasets. (3) Increasing model size improves robustness, although the gains diminish as models become larger. (4) Robustness to unseen corruptions can be improved through targeted data augmentation during training or speech enhancement at inference time. These findings highlight the importance of evaluating audio deepfake detectors under diverse real-world corruptions and developing more robust detection frameworks for practical deployment. We further advocate that future research on deepfake detection across all media should account for the diverse and unpredictable distortions encountered in real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。