arXiv:2506.09521eess.AScs.CL2025-06中稿 · INTERSPEECH 2025 u…被引 5

用语音文本内容就能破译匿名语音,暴露隐私漏洞

You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks

  • 用BERT模型分析语音文本,实现声纹识别攻击
  • 部分说话人识别错误率低至2%,平均错误率35%
  • 揭示数据集语义相似性对评测结果的干扰,适合安全研究者

说话人匿名系统在保留语言内容和情感的同时隐藏说话人身份。为评估其隐私保护效果,通常采用自动说话人验证(ASV)系统作为攻击手段。本研究通过将BERT语言模型用于ASV任务,评估了攻击者训练与测试数据集中说话人内部语言内容相似性的影响。在VoicePrivacy Attacker Challenge数据集上,该方法实现了平均等错误率(EER)35%,某些说话人甚至低至2%,仅依赖语音的文本内容。可解释性分析显示,系统决策与语句中的语义相似关键词相关,这源于LibriSpeech数据集的构建方式。研究建议重新设计VoicePrivacy数据集以实现公平、无偏的评测,并质疑当前以全局EER作为隐私评估标准的合理性。

原文摘要 · Abstract (English)

Speaker anonymization systems hide the identity of speakers while preserving other information such as linguistic content and emotions. To evaluate their privacy benefits, attacks in the form of automatic speaker verification (ASV) systems are employed. In this study, we assess the impact of intra-speaker linguistic content similarity in the attacker training and evaluation datasets, by adapting BERT, a language model, as an ASV system. On the VoicePrivacy Attacker Challenge datasets, our method achieves a mean equal error rate (EER) of 35%, with certain speakers attaining EERs as low as 2%, based solely on the textual content of their utterances. Our explainability study reveals that the system decisions are linked to semantically similar keywords within utterances, stemming from how LibriSpeech is curated. Our study suggests reworking the VoicePrivacy datasets to ensure a fair and unbiased evaluation and challenge the reliance on global EER for privacy evaluations.

语音隐私语言模型安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。