用微小噪声保护语音数据,防止被伪造成高质假声。
Mitigating Unauthorized Speech Synthesis for Voice Protection
- 在原始语音中加入不可察觉的扰动噪声,干扰语音合成模型学习。
- 保护后语音合成的模糊度评分从21.94%升至127.31%,显著降低伪造可能。
- 对降噪和数据增强有鲁棒性,适合保护个人语音隐私。
近年来,仅需少量语音样本即可完美复刻说话人声音,而恶意语音滥用(如电信诈骗)已对日常生活造成巨大威胁。因此,保护包含敏感信息(如个人声纹)的公开语音数据至关重要。以往防御方法主要针对音色相似性欺骗语音验证系统,但生成的深度伪造语音仍保持高质量。为此,我们提出一种高效、可迁移且鲁棒的主动防护技术——关键目标扰动(POP),在原始语音样本中加入不可察觉的最小误差噪声,使其难以被文本到语音(TTS)模型有效学习,从而阻止高质量深度伪造语音生成。我们在先进SOTA TTS模型上进行广泛实验,采用客观与主观评估指标全面验证该方法。结果表明,其在多种模型间具备优异效果与可迁移性:未受保护样本训练的语音合成器模糊度评分为21.94%,而经POP保护后的样本提升至127.31%。此外,该方法对降噪和数据增强具有强鲁棒性,大幅降低潜在风险。
原文摘要 · Abstract (English)
With just a few speech samples, it is possible to perfectly replicate a speaker's voice in recent years, while malicious voice exploitation (e.g., telecom fraud for illegal financial gain) has brought huge hazards in our daily lives. Therefore, it is crucial to protect publicly accessible speech data that contains sensitive information, such as personal voiceprints. Most previous defense methods have focused on spoofing speaker verification systems in timbre similarity but the synthesized deepfake speech is still of high quality. In response to the rising hazards, we devise an effective, transferable, and robust proactive protection technology named Pivotal Objective Perturbation (POP) that applies imperceptible error-minimizing noises on original speech samples to prevent them from being effectively learned for text-to-speech (TTS) synthesis models so that high-quality deepfake speeches cannot be generated. We conduct extensive experiments on state-of-the-art (SOTA) TTS models utilizing objective and subjective metrics to comprehensively evaluate our proposed method. The experimental results demonstrate outstanding effectiveness and transferability across various models. Compared to the speech unclarity score of 21.94% from voice synthesizers trained on samples without protection, POP-protected samples significantly increase it to 127.31%. Moreover, our method shows robustness against noise reduction and data augmentation techniques, thereby greatly reducing potential hazards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。