arXiv:2409.14919cs.SDeess.AS2024-09中稿 · publication in pro…被引 4

用对抗隐藏技术实现语音隐私保护,可控地改换说话人身份

Voice Conversion-based Privacy through Adversarial Information Hiding

  • 通过对抗信息隐藏控制语音身份信息泄露程度
  • 转换后语音身份感知被改变,但词汇内容保持完好
  • 比循环网络和语音合成流程更灵活且更安全

隐私保护型语音转换旨在仅移除承载身份信息的语音特征,同时保留其他语音特性。本文提出一种基于对抗信息隐藏的隐私保护语音转换机制,可控制身份相关信息的泄露程度,实现源语音特征保持与说话人身份修改之间的可控权衡。该方法优于未针对隐私设计的CycleGAN和StarGAN等语音转换技术,避免了身份信息不可控泄露的问题;也比ASR-TTS语音转换流程更灵活,后者会自然丢弃与文本相关的韵律信息。实验表明,所提系统能有效改变听者对说话人身份的感知,同时良好保持原始词汇内容。

原文摘要 · Abstract (English)

Privacy-preserving voice conversion aims to remove only the attributes of speech audio that convey identity information, keeping other speech characteristics intact. This paper presents a mechanism for privacy-preserving voice conversion that allows controlling the leakage of identity-bearing information using adversarial information hiding. This enables a deliberate trade-off between maintaining source-speech characteristics and modification of speaker identity. As such, the approach improves on voice-conversion techniques like CycleGAN and StarGAN, which were not designed for privacy, meaning that converted speech may leak personal information in unpredictable ways. Our approach is also more flexible than ASR-TTS voice conversion pipelines, which by design discard all prosodic information linked to textual content. Evaluations show that the proposed system successfully modifies perceived speaker identity whilst well maintaining source lexical content.

语音转换隐私保护对抗学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。