arXiv:2505.12332cs.SDcs.AI2025-05AAAI

针对扩散模型语音克隆,提出多维度防御框架,有效混淆身份并降低伪造质量。

VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning

  • 通过对抗扰动干扰参考音频,破坏扩散模型的生成过程。
  • 防御成功率高,显著降低语音克隆的可辨识性和感知质量。
  • 适合关注语音安全、防止恶意克隆的技术研发者使用。

扩散模型(DMs)在真实语音克隆(VC)中取得显著进展,但也增加了被恶意滥用的风险。现有针对传统语音克隆模型的主动防御方法因扩散模型复杂的生成机制而失效。为此,我们提出VoiceCloak,一个面向未经授权扩散模型语音克隆的多维主动防御框架,旨在混淆说话人身份并降低生成语音的感知质量。通过聚焦分析扩散模型中的特定漏洞,VoiceCloak在参考音频中引入对抗扰动,以干扰生成过程。具体而言,为混淆说话人身份,框架首先扭曲表示学习嵌入,最大化身份差异,遵循听觉感知原则;同时破坏关键的条件引导过程,尤其是注意力上下文,阻止声学特征对齐,从而削弱克隆效果。为实现降质目标,框架引入得分幅度放大,主动引导反向轨迹偏离高质量语音生成路径,并采用噪声引导语义破坏,干扰扩散模型对语音结构语义的捕捉,进一步降低输出质量。大量实验表明,VoiceCloak在抵御未经授权的扩散模型语音克隆方面具有优异的防御成功率。相关音频样本可访问 https://voice-cloak.github.io/VoiceCloak/。

原文摘要 · Abstract (English)

Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrupt the forgery process, but they have been proven incompatible with DMs due to the intricate generative mechanisms of diffusion. To bridge this gap, we introduce VoiceCloak, a multi-dimensional proactive defense framework with the goal of obfuscating speaker identity and degrading perceptual quality in potential unauthorized VC. To achieve these goals, we conduct a focused analysis to identify specific vulnerabilities within DMs, allowing VoiceCloak to disrupt the cloning process by introducing adversarial perturbations into the reference audio. Specifically, to obfuscate speaker identity, VoiceCloak first targets speaker identity by distorting representation learning embeddings to maximize identity variation, which is guided by auditory perception principles. Additionally, VoiceCloak disrupts crucial conditional guidance processes, particularly attention context, thereby preventing the alignment of vocal characteristics that are essential for achieving convincing cloning. Then, to address the second objective, VoiceCloak introduces score magnitude amplification to actively steer the reverse trajectory away from the generation of high-quality speech. Noise-guided semantic corruption is further employed to disrupt structural speech semantics captured by DMs, degrading output quality. Extensive experiments highlight VoiceCloak's outstanding defense success rate against unauthorized diffusion-based voice cloning. Audio samples of VoiceCloak are available at https://voice-cloak.github.io/VoiceCloak/.

语音安全扩散模型对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。