arXiv:2409.08913eess.AScs.LG2024-09中稿 · and presented at t…被引 21

提出混合方案,在保护情绪的同时有效隐藏说话人身份。

HLTCOE JHU Submission to the Voice Privacy Challenge 2024

  • 融合语音转换与文本转语音技术,动态混合生成隐私语音
  • 在半白盒攻击下实现超40%的错误识别率,保留47%的识别准确率
  • 适合需要平衡隐私与情感保真的语音匿名场景

我们提交了多项针对语音隐私挑战的系统,包括基于语音转换的方法(如kNN-VC和WavLM语音转换)以及基于文本转语音(TTS)的方法(如Whisper-VITS)。实验发现,语音转换方法虽能更好保留情绪内容,但在半白盒攻击场景下难以隐藏说话人身份;而TTS方法在匿名化方面表现更优,但情绪保留较差。为此,我们提出一种随机混合系统,旨在结合两类方法的优势,在保持47%的均匀准确率(UAR)的同时,实现超过40%的等错误率(EER),显著提升隐私保护能力。

原文摘要 · Abstract (English)

We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.

语音隐私匿名化语音转换TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。