arXiv:2507.12932cs.SDcs.MM2025-07中稿 · ACM MM 2025, Open-…被引 7

用少量语音数据生成通用频域噪声,实时保护音频隐私。

Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes

  • 基于黑盒知识与少量用户数据生成通用频域噪声。
  • 内存效率提升50至200倍,运行效率提升3至7000倍。
  • 适合需实时防护的语音应用,如通讯、语音助手。

语音深度伪造技术的快速发展引发了严重的音频隐私担忧,攻击者利用公开语音数据生成逼真的虚假音频,用于身份盗用、金融诈骗和误导性宣传。现有防御方法存在适应性差、难以扩展至长音频、依赖白盒信息及加密过程计算成本高等问题。为此,我们提出Enkidu,一种面向用户的隐私保护框架,通过少量用户数据和黑盒知识生成通用频域扰动。这些可调节的频域噪声块实现了实时、轻量级保护,具备跨长度音频的强泛化能力,并在保持听觉质量与语音可懂度的前提下,有效抵御语音深度伪造攻击。相比六种前沿反制方法,Enkidu在内存使用上降低50至200倍(最低仅0.004吉字节),运行时间效率提升3至7000倍(实时系数低至0.004)。在六种主流文本转语音模型和五种先进自动说话人验证模型上的实验表明,Enkidu在防御常规与自适应语音深度伪造攻击中具有有效性、可迁移性和实用性。代码已公开。

原文摘要 · Abstract (English)

The rapid advancement of voice deepfake technologies has raised serious concerns about user audio privacy, as attackers increasingly exploit publicly available voice data to generate convincing fake audio for malicious purposes such as identity theft, financial fraud, and misinformation campaigns. While existing defense methods offer partial protection, they face critical limitations, including weak adaptability to unseen user data, poor scalability to long audio, rigid reliance on white-box knowledge, and high computational and temporal costs during the encryption process. To address these challenges and defend against personalized voice deepfake threats, we propose Enkidu, a novel user-oriented privacy-preserving framework that leverages universal frequential perturbations generated through black-box knowledge and few-shot training on a small amount of user data. These highly malleable frequency-domain noise patches enable real-time, lightweight protection with strong generalization across variable-length audio and robust resistance to voice deepfake attacks, all while preserving perceptual quality and speech intelligibility. Notably, Enkidu achieves over 50 to 200 times processing memory efficiency (as low as 0.004 gigabytes) and 3 to 7000 times runtime efficiency (real-time coefficient as low as 0.004) compared to six state-of-the-art countermeasures. Extensive experiments across six mainstream text-to-speech models and five cutting-edge automated speaker verification models demonstrate the effectiveness, transferability, and practicality of Enkidu in defending against both vanilla and adaptive voice deepfake attacks. Our code is currently available.

音频隐私语音伪造实时防护频域扰动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。