arXiv:2511.16114cs.SD2025-11被引 1

用真实环境音保护语音训练,防克隆更可靠

SceneGuard: Training-Time Voice Protection with Scene-Consistent Audible Background Noise

  • 在语音中加入符合场景的可听背景音进行防护
  • 使语音克隆相似度下降5.5%,统计显著且保真度98.6%
  • 对压缩、降噪等常见反制手段依然有效,适合隐私保护场景

语音克隆技术可利用少量音频样本生成未经授权的语音,构成严重隐私威胁。现有基于不可感知对抗扰动的防御方法易受去噪、压缩等常见音频预处理影响。本文提出SceneGuard,一种训练时语音保护方法,通过在语音记录中添加与场景一致的可听背景噪声实现防护。不同于不可感知扰动,SceneGuard利用真实声学场景(如机场、街道、公园)生成语境恰当且抗干扰的保护噪声。在文本到语音训练攻击评估中,该方法实现5.5%的说话人相似度下降,具有极高的统计显著性(p < 10^{-15},Cohen's d = 2.18),同时保持98.6%的语音可懂度(STOI = 0.986)。鲁棒性测试表明,面对MP3压缩、谱减法、低通滤波和下采样等五种常见反制手段,SceneGuard仍能维持或增强保护效果。结果表明,可听且场景一致的噪声为训练时语音保护提供了比不可感知扰动更稳健的替代方案。源代码已开源:https://github.com/richael-sang/SceneGuard。

原文摘要 · Abstract (English)

Voice cloning technology poses significant privacy threats by enabling unauthorized speech synthesis from limited audio samples. Existing defenses based on imperceptible adversarial perturbations are vulnerable to common audio preprocessing such as denoising and compression. We propose SceneGuard, a training-time voice protection method that applies scene-consistent audible background noise to speech recordings. Unlike imperceptible perturbations, SceneGuard leverages naturally occurring acoustic scenes (e.g., airport, street, park) to create protective noise that is contextually appropriate and robust to countermeasures. We evaluate SceneGuard on text-to-speech training attacks, demonstrating 5.5% speaker similarity degradation with extremely high statistical significance (p < 10^{-15}, Cohen's d = 2.18) while preserving 98.6% speech intelligibility (STOI = 0.986). Robustness evaluation shows that SceneGuard maintains or enhances protection under five common countermeasures including MP3 compression, spectral subtraction, lowpass filtering, and downsampling. Our results suggest that audible, scene-consistent noise provides a more robust alternative to imperceptible perturbations for training-time voice protection. The source code are available at: https://github.com/richael-sang/SceneGuard.

语音保护隐私安全对抗防御场景噪声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。