arXiv:2505.17513cs.LGcs.CL2025-05EMNLP被引 6

语言微调可骗过语音反伪造系统,揭示新安全漏洞

What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection

  • 用文本级对抗攻击测试语音检测器脆弱性
  • 部分开源检测器准确率下降超60%,一商用系统从100%跌至32%
  • 适合关注语音安全与对抗防御的研究者

文本到语音技术的进展使逼真语音生成成为可能,催生了欺诈和冒名顶替等音频深度伪造攻击。尽管语音反欺骗系统至关重要,但以往研究主要关注声学层面的扰动,对语言变化的影响却未充分探索。本文通过引入文本级对抗攻击,评估开源与商业反欺骗检测器的语言敏感性。大量实验表明,即使轻微的语言扰动也能显著降低检测准确率:多个开源检测器-语音组合的攻击成功率超过60%;尤其一商用检测器在合成音频上的准确率从100%骤降至32%。通过特征归因分析,我们发现语言复杂度与模型级音频嵌入相似性均显著影响检测器脆弱性。进一步通过案例研究复现布拉德·皮特语音深度伪造诈骗事件,利用文本对抗攻击完全绕过商用检测系统。结果表明,需超越纯声学防御,将语言变异纳入鲁棒反欺骗系统设计。所有源代码将公开。

原文摘要 · Abstract (English)

Recent advances in text-to-speech technologies have enabled realistic voice generation, fueling audio-based deepfake attacks such as fraud and impersonation. While audio anti-spoofing systems are critical for detecting such threats, prior work has predominantly focused on acoustic-level perturbations, leaving the impact of linguistic variation largely unexplored. In this paper, we investigate the linguistic sensitivity of both open-source and commercial anti-spoofing detectors by introducing transcript-level adversarial attacks. Our extensive evaluation reveals that even minor linguistic perturbations can significantly degrade detection accuracy: attack success rates surpass 60% on several open-source detector-voice pairs, and notably one commercial detection accuracy drops from 100% on synthetic audio to just 32%. Through a comprehensive feature attribution analysis, we identify that both linguistic complexity and model-level audio embedding similarity contribute strongly to detector vulnerability. We further demonstrate the real-world risk via a case study replicating the Brad Pitt audio deepfake scam, using transcript adversarial attacks to completely bypass commercial detectors. These results highlight the need to move beyond purely acoustic defenses and account for linguistic variation in the design of robust anti-spoofing systems. All source code will be publicly available.

语音伪造对抗攻击安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。