arXiv:2410.04017eess.AS2024-10被引 2

对抗攻击让语音克隆变假声,新方法能有效防御。

Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System

  • 用对抗样本训练模型,提升抗干扰能力。
  • 用扩散模型还原被篡改的音频,恢复原声。
  • 适合关注语音安全与防伪造的研究者。

基于说话人嵌入的零样本语音合成系统可仅用少量数据为未见说话人生成高质量语音,但易受对抗攻击:攻击者在原始语音波形中添加人耳难以察觉的扰动,导致合成语音听起来像另一个人。此类漏洞带来严重的安全隐患,如说话人身份冒充和未经授权的语音篡改。本文研究两种主要防御策略:对抗训练与对抗净化。对抗训练通过在训练中引入对抗样本,增强模型对攻击的鲁棒性;对抗净化则利用扩散概率模型将被恶意扰动的音频还原为干净版本。实验表明,这些防御机制能显著降低对抗扰动的影响,提升零样本语音合成系统在对抗环境中的安全性与可靠性。

原文摘要 · Abstract (English)

Speaker embedding based zero-shot Text-to-Speech (TTS) systems enable high-quality speech synthesis for unseen speakers using minimal data. However, these systems are vulnerable to adversarial attacks, where an attacker introduces imperceptible perturbations to the original speaker's audio waveform, leading to synthesized speech sounds like another person. This vulnerability poses significant security risks, including speaker identity spoofing and unauthorized voice manipulation. This paper investigates two primary defense strategies to address these threats: adversarial training and adversarial purification. Adversarial training enhances the model's robustness by integrating adversarial examples during the training process, thereby improving resistance to such attacks. Adversarial purification, on the other hand, employs diffusion probabilistic models to revert adversarially perturbed audio to its clean form. Experimental results demonstrate that these defense mechanisms can significantly reduce the impact of adversarial perturbations, enhancing the security and reliability of speaker embedding based zero-shot TTS systems in adversarial environments.

语音合成对抗攻击安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。