arXiv:2504.09839cs.SDcs.AI2025-04中稿 · USENIX Security 20…被引 12

通过隐形扰动保护语音,防止深度伪造滥用

SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis

论文配图:SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis
图 1 · 摘自论文原文
  • 用代理模型生成通用扰动,嵌入语音防克隆
  • 在真实测试中实现高鲁棒性与实时处理能力
  • 适合需防范语音伪造的个人与机构使用

语音合成技术带来便利,但逼真的深度伪造音频已引发严重风险。恶意攻击者可未经授权收集他人语音并克隆声音用于非法活动(如电信诈骗)。现有防御方法难以有效阻止深度伪造,且易受鲁棒训练技术攻击。为此,我们提出安全语音防护框架 SafeSpeech,通过在上传前对原始语音嵌入不可感知的扰动,防止高质量语音合成。SafeSpeech 设计了鲁棒且通用的主动防护技术——语音扰动隐藏(SPEC),利用代理模型生成适用于各类生成式模型的通用扰动,并优化了时间与频域中的人耳感知效果。我们在多个先进模型与数据集上进行了全面实验,涵盖主观与客观评估。结果表明,SafeSpeech 在语音保护有效性、迁移性及对抗高级自适应攻击方面均达到当前最优水平,且具备实际部署所需的实时性能。源代码已公开于 https://github.com/wxzyd123/SafeSpeech。

原文摘要 · Abstract (English)

Speech synthesis technology has brought great convenience, while the widespread usage of realistic deepfake audio has triggered hazards. Malicious adversaries may unauthorizedly collect victims' speeches and clone a similar voice for illegal exploitation (\textit{e.g.}, telecom fraud). However, the existing defense methods cannot effectively prevent deepfake exploitation and are vulnerable to robust training techniques. Therefore, a more effective and robust data protection method is urgently needed. In response, we propose a defensive framework, \textit{\textbf{SafeSpeech}}, which protects the users' audio before uploading by embedding imperceptible perturbations on original speeches to prevent high-quality synthetic speech. In SafeSpeech, we devise a robust and universal proactive protection technique, \textbf{S}peech \textbf{PE}rturbative \textbf{C}oncealment (\textbf{SPEC}), that leverages a surrogate model to generate universally applicable perturbation for generative synthetic models. Moreover, we optimize the human perception of embedded perturbation in terms of time and frequency domains. To evaluate our method comprehensively, we conduct extensive experiments across advanced models and datasets, both subjectively and objectively. Our experimental results demonstrate that SafeSpeech achieves state-of-the-art (SOTA) voice protection effectiveness and transferability and is highly robust against advanced adaptive adversaries. Moreover, SafeSpeech has real-time capability in real-world tests. The source code is available at \href{https://github.com/wxzyd123/SafeSpeech}{https://github.com/wxzyd123/SafeSpeech}.

语音安全深度伪造隐私保护对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。