分析语音匿名化中个体重识别风险差异,发现风险由多方交互决定。
A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization
- 按说话人逐个评估匿名化风险,使用最坏情况下的链接性指标。
- 近5000名说话人中,重识别难度分布极不均衡,无统一易识别人群。
- 风险受攻击者、匿名系统和语音时长共同影响,需针对性评估。
语音匿名化通常采用均值指标(如等错误率)进行评估,可能掩盖个体间的重识别风险差异。本文基于链接性指标,在最坏情形下对近5000名说话人开展大规模逐说话人隐私分析,覆盖多种匿名化系统、攻击者架构及对话长度。结果表明,链接性得分在说话人层面高度极化,易识别与难识别的说话人集合随配置显著变化。单一因素无法解释说话人脆弱性,重识别风险源于攻击者、匿名化器与可用语音量之间的相互作用。研究挑战了固有说话人隐私风险的观念,强调评估协议必须显式考虑攻击者与匿名化器的组合。
原文摘要 · Abstract (English)
Speech anonymization is commonly evaluated using averagecase metrics such as the equal error rate, which can hide large disparities in re-identification risks across individuals. In this paper, we conduct a large-scale per-speaker privacy analysis using a linkability-based metric under a worst-case scenario. Nearly 5,000 speakers are evaluated across multiple anonymization systems, attacker architectures, and conversation lengths. While linkability scores are highly polarized at the speaker level, the sets of easy to re-identify and hard to re-identify speakers vary substantially across configurations. We show that no single factor explains speaker vulnerability. Instead, the re-identification risk emerges from the interaction between the attacker, the anonymizer, and the amount of available speech. These results challenge the notion of intrinsic speaker-level privacy risks and emphasize the need for evaluation protocols that are explicitly conditioned on the attacker and anonymizer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。