arXiv:2601.02444cs.SDcs.AI2026-01

提出新净化框架,可有效破解语音防克隆扰动。

VocalBridge: Latent Diffusion-Bridge Purification for Defeating Perturbation-Based Voiceprint Defenses

  • 在EnCodec隐空间构建扩散桥,无须文本即可净化带扰动语音。
  • 对防克隆语音的还原成功率超现有方法,逼近真实语音质量。
  • 适合研究语音安全、防御机制的学者与工程师参考。

语音合成技术的快速发展加剧了语音克隆带来的安全与隐私风险。近期防御方案通过在语音中嵌入保护性扰动来隐藏说话人身份,同时保持可理解性。然而,攻击者可利用先进净化技术去除这些扰动,恢复原始声学特征并生成可克隆语音。现有净化方法多针对自动语音识别中的对抗噪声,无法有效抑制定义说话人身份的细微声学线索,对说话人验证攻击(SVA)效果不佳。为此,我们提出扩散桥(VocalBridge),在EnCodec隐空间学习从扰动到清晰语音的映射。采用带时间条件的1D U-Net与余弦噪声调度,实现无需语音转录的高效净化,并保留说话人判别结构。进一步引入基于Whisper的音素变体,以轻量级时序引导增强效果。实验表明,该方法在从受保护语音中恢复可克隆语音方面显著优于现有净化手段。结果揭示了当前扰动防御的脆弱性,凸显应对演进的语音克隆与说话人验证威胁需更强防护机制。

原文摘要 · Abstract (English)

The rapid advancement of speech synthesis technologies, including text-to-speech (TTS) and voice conversion (VC), has intensified security and privacy concerns related to voice cloning. Recent defenses attempt to prevent unauthorized cloning by embedding protective perturbations into speech to obscure speaker identity while maintaining intelligibility. However, adversaries can apply advanced purification techniques to remove these perturbations, recover authentic acoustic characteristics, and regenerate cloneable voices. Despite the growing realism of such attacks, the robustness of existing defenses under adaptive purification remains insufficiently studied. Most existing purification methods are designed to counter adversarial noise in automatic speech recognition (ASR) systems rather than speaker verification or voice cloning pipelines. As a result, they fail to suppress the fine-grained acoustic cues that define speaker identity and are often ineffective against speaker verification attacks (SVA). To address these limitations, we propose Diffusion-Bridge (VocalBridge), a purification framework that learns a latent mapping from perturbed to clean speech in the EnCodec latent space. Using a time-conditioned 1D U-Net with a cosine noise schedule, the model enables efficient, transcript-free purification while preserving speaker-discriminative structure. We further introduce a Whisper-guided phoneme variant that incorporates lightweight temporal guidance without requiring ground-truth transcripts. Experimental results show that our approach consistently outperforms existing purification methods in recovering cloneable voices from protected speech. Our findings demonstrate the fragility of current perturbation-based defenses and highlight the need for more robust protection mechanisms against evolving voice-cloning and speaker verification threats.

语音克隆扩散模型语音安全防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。