arXiv:2505.19951cs.SDcs.AI2025-05被引 3

用新损失函数提升语音隐私保护的隐蔽性和兼容性

Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy

  • 引入指数总变差损失,增强对抗补丁的隐蔽性
  • 在不同音频长度下均保持高效语音匿名化效果
  • 适合需要长期保护语音身份的智能设备场景

当前深度学习语音模型虽广泛应用,但个人数据(如身份和语义内容)的安全处理仍存疑。为防止恶意识别,提出语音匿名化方法。现有基于通用对抗补丁(UAP)的方法存在音频质量下降、语音识别性能降低、跨模型迁移能力弱及输入长度依赖等问题。本文提出新型指数总变差(Exponential Total Variance, TV)损失函数,并实验证明其可有效提升UAP的强度与不可察觉性。同时设计一种可扩展的UAP插入机制,在多种音频长度下均表现稳定,显著改善迁移能力与鲁棒性。

原文摘要 · Abstract (English)

Deep learning voice models are commonly used nowadays, but the safety processing of personal data, such as human identity and speech content, remains suspicious. To prevent malicious user identification, speaker anonymization methods were proposed. Current methods, particularly based on universal adversarial patch (UAP) applications, have drawbacks such as significant degradation of audio quality, decreased speech recognition quality, low transferability across different voice biometrics models, and performance dependence on the input audio length. To mitigate these drawbacks, in this work, we introduce and leverage the novel Exponential Total Variance (TV) loss function and provide experimental evidence that it positively affects UAP strength and imperceptibility. Moreover, we present a novel scalable UAP insertion procedure and demonstrate its uniformly high performance for various audio lengths.

语音隐私对抗补丁匿名化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。