通过语音时序动态提取说话人特征,提升语音匿名系统攻击效果。
Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems
- 从语音节奏、语调等时序特征中提取上下文相关的时长嵌入表示。
- 在原始和匿名语音上均显著提升语音验证性能,优于已有方法。
- 揭示了语音验证与匿名系统中的潜在漏洞,适合安全研究者参考。
语音的时序动态特性,包括节奏、语调和语速的变化,蕴含着关于说话人身份的重要且独特的信息。本文提出一种新方法,通过从语音时序动态中提取上下文依赖的时长嵌入来表征说话人特征。我们构建了基于这些表示的新攻击模型,并分析了语音验证与语音匿名化系统中的潜在漏洞。实验结果表明,与文献中报道的更简单时序动态表示相比,所提出的攻击模型在原始数据和匿名化数据上的语音验证性能均有显著提升。
原文摘要 · Abstract (English)
The temporal dynamics of speech, encompassing variations in rhythm, intonation, and speaking rate, contain important and unique information about speaker identity. This paper proposes a new method for representing speaker characteristics by extracting context-dependent duration embeddings from speech temporal dynamics. We develop novel attack models using these representations and analyze the potential vulnerabilities in speaker verification and voice anonymization systems.The experimental results show that the developed attack models provide a significant improvement in speaker verification performance for both original and anonymized data in comparison with simpler representations of speech temporal dynamics reported in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。