研究深度伪造语音检测器如何被恶意攻击,揭示现有方法的脆弱性。
Can DeepFake Speech be Reliably Detected?
- 设计白盒与黑盒攻击,测试检测器在伪装下的失效情况
- 攻击后检测准确率下降超过40%,且人类难辨真伪
- 提醒需开发更鲁棒的检测技术,尤其面对主动对抗攻击
近年来,具备语音克隆能力的文本转语音(TTS)系统快速发展,使语音模仿变得容易,引发虚假信息传播和欺诈等伦理与法律问题。尽管已有合成语音检测器(SSD)用于应对,但其在音频经编码转换、播放或背景噪声处理后性能显著下降,即存在“测试域偏移”问题。更严重的是,攻击者可主动篡改合成语音以欺骗检测器。本文首次系统研究针对先进开源SSD的主动恶意攻击,从攻击有效性与隐蔽性角度,通过硬指标与人工评分评估白盒、黑盒攻击及其迁移能力。结果表明,当前检测方法面临严峻挑战,亟需提升抗对抗攻击的鲁棒性。
原文摘要 · Abstract (English)
Recent advances in text-to-speech (TTS) systems, particularly those with voice cloning capabilities, have made voice impersonation readily accessible, raising ethical and legal concerns due to potential misuse for malicious activities like misinformation campaigns and fraud. While synthetic speech detectors (SSDs) exist to combat this, they are vulnerable to ``test domain shift", exhibiting decreased performance when audio is altered through transcoding, playback, or background noise. This vulnerability is further exacerbated by deliberate manipulation of synthetic speech aimed at deceiving detectors. This work presents the first systematic study of such active malicious attacks against state-of-the-art open-source SSDs. White-box attacks, black-box attacks, and their transferability are studied from both attack effectiveness and stealthiness, using both hardcoded metrics and human ratings. The results highlight the urgent need for more robust detection methods in the face of evolving adversarial threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。