通过可解释组件增强语音转换模型,提升说话人匿名化隐私保护效果
Private kNN-VC: Interpretable Anonymization of Converted Speech
- 引入可解释模块,分别处理音素时长和变化特征以增强匿名性
- 隐私保护能力显著提升,证明韵律特征泄露说话人身份
- 适用于关注语音隐私安全与可解释性的研究者
说话人匿名化旨在隐藏说话人身份的同时保留语音可用性。现有评估多依赖在匿名化语音上训练的说话人识别模型,虽为强攻击方式,但难以判断具体哪些语音特征被用于识别。本研究旨在揭示这些关键特征。基于性能较差的语音转换模型kNN-VC,我们引入两个可解释组件,分别对音素时长和变化特征进行匿名化处理。实验表明,该改进显著提升隐私保护效果,证实所研究的韵律特征确实编码了说话人身份,并被隐私攻击利用。此外,目标选择算法的调整对隐私攻击结果有显著影响。
原文摘要 · Abstract (English)
Speaker anonymization seeks to conceal a speaker's identity while preserving the utility of their speech. The achieved privacy is commonly evaluated with a speaker recognition model trained on anonymized speech. Although this represents a strong attack, it is unclear which aspects of speech are exploited to identify the speakers. Our research sets out to unveil these aspects. It starts with kNN-VC, a powerful voice conversion model that performs poorly as an anonymization system, presumably because of prosody leakage. To test this hypothesis, we extend kNN-VC with two interpretable components that anonymize the duration and variation of phones. These components increase privacy significantly, proving that the studied prosodic factors encode speaker identity and are exploited by the privacy attack. Additionally, we show that changes in the target selection algorithm considerably influence the outcome of the privacy attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。