语音匿名化隐私保护可能被高估,研究揭示风险并提供检测方法。
The Risks and Detection of Overestimated Privacy Protection in Voice Anonymisation
- 用说话人验证系统评估匿名化效果时,模型训练不足会夸大隐私保护能力。
- 最严重情况下隐私保护被高估74%相对值,文献中存在类似误判实例。
- 提出可信赖性检测机制,已开源,适用于语音隐私评测工具链。
语音匿名化旨在隐藏语音记录中的说话人身份。隐私保护程度通常通过匿名化后使用说话人验证系统重新识别的难度来估计。因此,评估结果依赖于验证模型与匿名化系统的性能。当验证模型训练不佳或数据不匹配时,可能存在隐私保护被过度估计的风险。本文揭示了这种隐蔽风险,并展示了文献中夸大性能的实例。在最严重情况下,性能被高估了74%相对值。随后,我们提出一种检测评估不可靠性的方法,并证明其能识别所有文中展示的过度假设场景。该解决方案已作为2024年语音隐私挑战赛评估工具包的开源分支发布。
原文摘要 · Abstract (English)
Voice anonymisation aims to conceal the voice identity of speakers in speech recordings. Privacy protection is usually estimated from the difficulty of using a speaker verification system to re-identify the speaker post-anonymisation. Performance assessments are therefore dependent on the verification model as well as the anonymisation system. There is hence potential for privacy protection to be overestimated when the verification system is poorly trained, perhaps with mismatched data. In this paper, we demonstrate the insidious risk of overestimating anonymisation performance and show examples of exaggerated performance reported in the literature. For the worst case we identified, performance is overestimated by 74% relative. We then introduce a means to detect when performance assessment might be untrustworthy and show that it can identify all overestimation scenarios presented in the paper. Our solution is openly available as a fork of the 2024 VoicePrivacy Challenge evaluation toolkit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。