提出新评估框架,揭示语音匿名化在真实攻击下的局部隐私泄露风险
Goodbye Equal Error Rate, Hello Local Information Disclosure: Evaluating Voice Anonymisation against 1-to-N Linkage Threats
- 基于信息论构建1对N链接攻击下的隐私评估框架
- 发现顶尖系统在近完美EER下仍存在每语音片段最高1比特泄露
- 适合关注隐私安全合规的开发者与监管者参考
语音匿名化旨在保护说话人身份。当前其隐私评估高度依赖等错误率(EER),该指标原为生物识别验证设计,全局聚合得分,隐含假设攻击者仅进行1对1匹配判断。这与真实世界中的数据库链接攻击(1对N封闭集搜索)存在威胁模型错配,导致局部隐私失败被全局平均掩盖。尽管近期已有1对N指标解决聚合问题,但未反映生物特征证据强度。本文提出一个模块化、信息论驱动的评估框架,针对1对N链接威胁建模。核心指标局部信息泄露(LID)通过将原始相似度分数校准为攻击者对每个注册身份的后验置信度,量化单次语音试验的隐私损失比特数。对VoicePrivacy 2024挑战赛中表现最佳系统的评估显示,尽管EER接近完美(48%),仍存在局部漏洞,最坏情况下每试验泄露达1比特(使攻击成功率翻倍于随机猜测)。结果表明,采用局部隐私度量对捕捉最坏情况风险及符合严格隐私法规至关重要。
原文摘要 · Abstract (English)
Voice anonymisation aims to protect speaker identity. Currently, its empirical privacy evaluation heavily relies on the Equal Error Rate (EER). Originally designed for biometric verification, EER aggregates scores globally, implicitly assuming an attacker is only trying to verify if two specific voice samples match (a 1-to-1 comparison). This introduces a threat model mismatch with real-world database linkage attacks, where an attacker searches across a fixed set of N enrolled identities (a 1-to-N closed-set search), allowing global averages to obscure localised privacy failures. While recent 1-to-N metrics address this aggregation issue, they abstract away the magnitude of the biometric evidence. In this paper, we propose a modular, information-theoretic evaluation framework explicitly designed for the 1-to-N linkage threat model. Within this framework, our core metric, Local Information Disclosure (LID), quantifies the exact privacy loss of a single trial utterance in bits by calibrating its raw similarity scores into the attacker's posterior confidence for each enrolled identity. Evaluating top-performing systems from the VoicePrivacy 2024 Challenge reveals that systems exhibiting near-perfect EERs (48 %) can still suffer from localised vulnerabilities with worst-case disclosures reaching 1 bit per trial utterance (effectively doubling the attacker's success rate over a random guess). We demonstrate that adopting localised privacy metrics is essential for capturing worst-case risks and aligning with strict privacy regulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。