arXiv:2508.17660cs.SDcs.CR2025-08中稿 · AsiaCCS 2025被引 3

ClearMask通过频域滤波和声学风格迁移,无损防御语音深度伪造攻击。

ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks

  • 在频谱图上选择性滤除频率,不加噪但破坏生成模型特征
  • 保护后语音对人耳和语音识别模型均保持自然,误判率低于5%
  • 适合实时会议、语音消息等场景的语音隐私防护

语音深度伪造攻击通过合成逼真语音实施恶意行为,已成为严重威胁。现有防御方法通常向语音注入噪声以干扰语音编码器,但会降低音质且需预先了解攻击方式,适用场景有限。尤其在虚拟会议、语音消息等实时音频中,仍面临安全风险。为此,我们提出ClearMask,一种无噪且保留自然度的语音防御机制。不同于传统方法,ClearMask通过选择性过滤音频梅尔频谱图中的特定频率,诱导可转移的语音特征损失,而无需添加噪声。随后采用音频风格迁移进一步迷惑语音解码器,同时保持听觉自然度。最后引入优化混响,干扰语音生成模型输出,而不影响语音自然性。此外,我们开发了LiveMask,通过通用频率滤波器与混响生成器实现实时流式语音防护。实验表明,ClearMask和LiveMask能有效防止未见过的语音合成模型及黑盒API服务欺骗说话人验证模型与人类听者。同时,ClearMask对试图从保护语音中恢复原始信号的自适应攻击也表现出强鲁棒性。

原文摘要 · Abstract (English)

Voice deepfake attacks, which artificially impersonate human speech for malicious purposes, have emerged as a severe threat. Existing defenses typically inject noise into human speech to compromise voice encoders in speech synthesis models. However, these methods degrade audio quality and require prior knowledge of the attack approaches, limiting their effectiveness in diverse scenarios. Moreover, real-time audios, such as speech in virtual meetings and voice messages, are still exposed to voice deepfake threats. To overcome these limitations, we propose ClearMask, a noise-free defense mechanism against voice deepfake attacks. Unlike traditional approaches, ClearMask modifies the audio mel-spectrogram by selectively filtering certain frequencies, inducing a transferable voice feature loss without injecting noise. We then apply audio style transfer to further deceive voice decoders while preserving perceived sound quality. Finally, optimized reverberation is introduced to disrupt the output of voice generation models without affecting the naturalness of the speech. Additionally, we develop LiveMask to protect streaming speech in real-time through a universal frequency filter and reverberation generator. Our experimental results show that ClearMask and LiveMask effectively prevent voice deepfake attacks from deceiving speaker verification models and human listeners, even for unseen voice synthesis models and black-box API services. Furthermore, ClearMask demonstrates resilience against adaptive attackers who attempt to recover the original audio signal from the protected speech samples.

语音防御深度伪造实时防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。