通过重组语音片段增强攻击者识别隐匿语音中说话人信息的能力
SegReConcat: A Data Augmentation Method for Voice Anonymization Attack
- 将语音按词分段后随机或按相似性重排并拼接,破坏长期语境线索
- 在7个匿名化系统中,有5个被成功攻破,提升识别率
- 适用于评估语音匿名化安全性的攻击方研究,尤其关注隐私泄露风险
语音匿名化旨在隐藏说话人身份的同时保持语音数据可用性。然而,残留的说话人特征常导致隐私泄露风险。本文提出一种攻击端数据增强方法 SegReConcat,用于提升自动说话人验证系统的攻击能力。该方法在词级别对匿名语音进行分段,采用随机或基于相似性的策略重新排列片段,并与原始语音拼接,使攻击者能从多角度学习源说话人特征。在 VoicePrivacy Attacker Challenge 2024 框架下,针对七种语音匿名化系统进行评估,SegReConcat 在其中五种系统上实现了去匿名化性能提升。
原文摘要 · Abstract (English)
Anonymization of voice seeks to conceal the identity of the speaker while maintaining the utility of speech data. However, residual speaker cues often persist, which pose privacy risks. We propose SegReConcat, a data augmentation method for attacker-side enhancement of automatic speaker verification systems. SegReConcat segments anonymized speech at the word level, rearranges segments using random or similarity-based strategies to disrupt long-term contextual cues, and concatenates them with the original utterance, allowing an attacker to learn source speaker traits from multiple perspectives. The proposed method has been evaluated in the VoicePrivacy Attacker Challenge 2024 framework across seven anonymization systems, SegReConcat improves de-anonymization on five out of seven systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。