arXiv:2505.14862cs.SDcs.AI2025-05被引 18

用录音重放攻击让语音伪造检测失效,暴露现有系统漏洞。

Replay Attacks Against Audio Deepfake Detection

  • 通过不同人和麦克风重录伪造语音,欺骗检测模型
  • 顶级模型错误率从4.7%升至18.2%,性能大幅下降
  • 适合语音安全与对抗攻防研究者关注

我们揭示了录音重放攻击如何破坏语音伪造检测:通过在不同说话人和麦克风上播放并重新录制伪造语音,使样本看似真实。为深入研究该现象,我们构建了ReplayDF数据集,源自M-AILABS和MLAAD,涵盖六种语言、四种TTS模型及109种话者-麦克风组合,包含多样声学条件,部分极具挑战性。对六个开源检测模型在五个数据集上的分析显示显著脆弱性——表现最佳的W2V2-AASIST模型等错误率(EER)从4.7%飙升至18.2%。即使采用自适应房间冲激响应(RIR)再训练,性能仍受限,EER维持在11.0%。相关数据集已开放用于非商业研究。

原文摘要 · Abstract (English)

We show how replay attacks undermine audio deepfake detection: By playing and re-recording deepfake audio through various speakers and microphones, we make spoofed samples appear authentic to the detection model. To study this phenomenon in more detail, we introduce ReplayDF, a dataset of recordings derived from M-AILABS and MLAAD, featuring 109 speaker-microphone combinations across six languages and four TTS models. It includes diverse acoustic conditions, some highly challenging for detection. Our analysis of six open-source detection models across five datasets reveals significant vulnerability, with the top-performing W2V2-AASIST model's Equal Error Rate (EER) surging from 4.7% to 18.2%. Even with adaptive Room Impulse Response (RIR) retraining, performance remains compromised with an 11.0% EER. We release ReplayDF for non-commercial research use.

语音伪造对抗攻击检测漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。