LibriSpeech数据集存在内容泄露,导致说话人身份可被词汇识别
Content Leakage in LibriSpeech and Its Impact on the Privacy Evaluation of Speaker Anonymization
- 利用说话人阅读书籍的词汇特征进行身份识别
- 在LibriSpeech上识别率超90%,表明隐私评估存在漏洞
- 推荐使用EdAcc数据集提升匿名化评测可靠性
说话人匿名化旨在隐藏说话人身份,但不考虑语言内容。本研究揭示了常用评测数据集LibriSpeech的一个缺陷:不同说话人所读书籍差异显著,导致可通过词汇特征识别说话人身份,即使使用完美的匿名化方法也无法防止此类信息泄露。相比之下,EdAcc数据集中的词汇区分度较低,仅少数说话人可被词汇识别,迫使攻击者转向其他线索,更真实反映匿名化系统的安全性。此外,EdAcc包含更多自发性口语和多样化说话人,与LibriSpeech形成互补,为评估匿名化模型提供更全面视角。
原文摘要 · Abstract (English)
Speaker anonymization aims to conceal a speaker's identity, without considering the linguistic content. In this study, we reveal a weakness of Librispeech, the dataset that is commonly used to evaluate anonymizers: the books read by the Librispeech speakers are so distinct, that speakers can be identified by their vocabularies. Even perfect anonymizers cannot prevent this identity leakage. The EdAcc dataset is better in this regard: only a few speakers can be identified through their vocabularies, encouraging the attacker to look elsewhere for the identities of the anonymized speakers. EdAcc also comprises spontaneous speech and more diverse speakers, complementing Librispeech and giving more insights into how anonymizers work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。