语音时序特征暴露说话人身份,影响声纹验证与匿名化效果
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
- 仅用音素持续时间做声纹验证,无需其他语音特征
- 原始和匿名语音中的音素时长均能泄露说话人信息
- 强调需调整发音时长特征以提升语音匿名保护能力
本文研究语音时序动态在自动说话人验证与说话人语音匿名化任务中的影响。提出基于音素持续时间的若干度量方法,用于自动说话人验证。实验结果表明,音素持续时间会泄露部分说话人信息,即使在原始语音和匿名化语音中也能识别说话人身份。因此,本工作强调必须考虑说话人的语速及音素持续时间特征,并对其进行调整,以构建具备强隐私保护能力的匿名化系统。
原文摘要 · Abstract (English)
In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verification based only on phoneme durations. Experimental results demonstrate that phoneme durations leak some speaker information and can reveal speaker identity from both original and anonymized speech. Thus, this work emphasizes the importance of taking into account the speaker's speech rate and, more importantly, the speaker's phonetic duration characteristics, as well as the need to modify them in order to develop anonymization systems with strong privacy protection capacity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。