提出新指标,量化语音匿名系统泄露软生物特征的风险
Measuring Soft Biometric Leakage in Speaker De-Identification Systems
- 设计统一评分体系SBLS,检测语音匿名后仍可推断的非唯一特征
- 五种主流系统均暴露显著漏洞,零样本攻击可精准还原年龄、性别等信息
- 适合关注隐私安全与语音匿名技术评估的研究者参考
我们使用重识别指从匿名语音输出中恢复原始说话人身份的过程。语音去标识系统旨在降低重识别风险,但现有评估多局限于个体层面指标,忽视了软生物特征泄露的广泛风险。本文提出软生物特征泄露评分(SBLS),一种统一方法,用于量化对非唯一特征(如信道类型、年龄范围、方言、性别、说话风格)的零样本推理攻击的抵抗能力。SBLS整合三个要素:使用预训练分类器进行直接属性推断、通过互信息分析检测关联性、跨交叉属性的子群体鲁棒性测试。在公开预训练模型下应用SBLS,发现所有五种评估的去标识系统均存在显著脆弱性。结果表明,仅使用预训练模型而无需原始语音或系统细节,攻击者仍可可靠恢复软生物特征信息,暴露出标准分布度量无法捕捉的根本性缺陷。
原文摘要 · Abstract (English)
We use the term re-identification to refer to the process of recovering the original speaker's identity from anonymized speech outputs. Speaker de-identification systems aim to reduce the risk of re-identification, but most evaluations focus only on individual-level measures and overlook broader risks from soft biometric leakage. We introduce the Soft Biometric Leakage Score (SBLS), a unified method that quantifies resistance to zero-shot inference attacks on non-unique traits such as channel type, age range, dialect, sex of the speaker, or speaking style. SBLS integrates three elements: direct attribute inference using pre-trained classifiers, linkage detection via mutual information analysis, and subgroup robustness across intersecting attributes. Applying SBLS with publicly available classifiers, we show that all five evaluated de-identification systems exhibit significant vulnerabilities. Our results indicate that adversaries using only pre-trained models - without access to original speech or system details - can still reliably recover soft biometric information from anonymized output, exposing fundamental weaknesses that standard distributional metrics fail to capture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。