破解语音转换伪装,还原被篡改音频中的真实说话人身份
Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion

- 设计三分支框架,分别识别转换机制与目标说话人特征
- 在7种主流语音转换方法上达到90.99%的还原准确率
- 适用于电话信道、未知语言等复杂场景,适合司法鉴伪应用
语音转换(VC)对生物识别安全构成重大威胁,使攻击者能冒充目标说话人。在司法取证中,从转换后的音频中恢复原始说话人身份对缩小嫌疑人范围至关重要。为此,我们提出TRIDENT,一种用于重建源说话人身份的追溯框架。该框架采用三重结构:主提取器及两个辅助分支。第一个辅助分支识别潜在的语音转换机制,利用攻击者通常使用主流模型变体的特性;第二个辅助分支提取目标说话人的隐式表征,帮助分离复合音频中的目标特异性特征。主提取器结合两个辅助分支的信息,解耦干扰因素,提炼出高度判别性的源说话人身份表征。实验表明,TRIDENT在对抗7种先进语音转换方法时,准确率达90.99%。此外,其在电话信道、未见语言和自适应场景下仍保持鲁棒性能。
原文摘要 · Abstract (English)
Voice conversion (VC) poses a significant threat to biometric security by allowing attackers to impersonate target speakers. In forensic contexts, recovering the source speaker's identity from converted audio is vital for narrowing the field of suspects. To address this, we propose TRIDENT, a retracing framework designed to restore a source speaker's original identity from a converted audio sample. TRIDENT utilizes a three-pronged architecture consisting of a primary extractor and two auxiliary branches. The first auxiliary branch identifies the underlying voice conversion mechanism. This design acknowledges that even if the exact conversion strategy is unknown, a high-performance model adopted by the attacker is typically a derivative or variant of established mainstream ones. The second auxiliary branch extracts a latent representation of the target speaker, facilitating the isolation of target-specific traits from the composite converted audio sample. Finally, the main extractor leverages insights from both auxiliary branches to decouple confounding factors and distill a highly discriminative representation of the source speaker's identity. Experimental results demonstrate that TRIDENT achieves an accuracy as high as 90.99% against 7 state-of-the-art voice conversion methods. Furthermore, TRIDENT maintains robust performance under challenging conditions, including telephony channels, unseen languages, and adaptive scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。