arXiv:2409.14712eess.AScs.SD2024-09被引 5

用房间冲激响应伪造语音,可让深度伪造音频绕过检测系统。

Room Impulse Responses help attackers to evade Deep Fake Detection

  • 通过模拟房间声学环境增强伪造语音
  • 使现有最优检测系统的错误率翻倍至5.16%
  • 适合研究语音安全与对抗攻击的学者

ASVspoof 2021基准测试是广泛使用的反欺骗评估框架,包含逻辑访问(LA)和深度伪造(DF)两个子集,涵盖不同编码特征与压缩伪影的样本。当前最先进(SOTA)系统在LA子集上达到0.87%的等错误率(EER),在DF子集上为2.58%。然而,基准测试准确率无法保证真实场景下的鲁棒性。本文研究利用房间冲激响应(RIRs)增强伪造语音以提高其逃避检测的可能性。结果表明,该简单方法显著提升逃逸率,使SOTA系统的EER翻倍。为应对此类攻击,我们使用大规模合成/仿真RIR数据集扩充训练数据,结果显示对混响伪造语音和原始样本均有显著改进,将DF任务的EER降至2.13%。

原文摘要 · Abstract (English)

The ASVspoof 2021 benchmark, a widely-used evaluation framework for anti-spoofing, consists of two subsets: Logical Access (LA) and Deepfake (DF), featuring samples with varied coding characteristics and compression artifacts. Notably, the current state-of-the-art (SOTA) system boasts impressive performance, achieving an Equal Error Rate (EER) of 0.87% on the LA subset and 2.58% on the DF. However, benchmark accuracy is no guarantee of robustness in real-world scenarios. This paper investigates the effectiveness of utilizing room impulse responses (RIRs) to enhance fake speech and increase their likelihood of evading fake speech detection systems. Our findings reveal that this simple approach significantly improves the evasion rate, doubling the SOTA system's EER. To counter this type of attack, We augmented training data with a large-scale synthetic/simulated RIR dataset. The results demonstrate significant improvement on both reverberated fake speech and original samples, reducing DF task EER to 2.13%.

语音伪造对抗攻击声学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。