评测语音去标识系统泄露身份风险,发现现有技术仍存严重隐私隐患。
Evaluating Identity Leakage in Speaker De-Identification Systems
- 构建三重误差率基准,量化语音去标识后残留的身份泄露。
- 最优系统表现仅略高于随机猜测,最差系统前50名命中率达45%。
- 揭示当前语音去标识技术存在持续隐私风险,适合关注语音安全的研究者。
语音去标识旨在隐藏说话人身份的同时保持语音可理解性。本文引入一个基准,通过三种互补的错误率来量化残留身份泄露:等错误率、累计匹配特征命中率,以及通过典型相关分析和普鲁斯特分析测量的嵌入空间相似性。评估结果显示,所有最先进的语音去标识系统均存在身份信息泄露。评估中表现最好的系统仅略优于随机猜测,而表现最差的系统在前50个候选者中的命中率高达45%(基于CMC)。这些发现凸显了当前语音去标识技术中持续存在的隐私风险。
原文摘要 · Abstract (English)
Speaker de-identification aims to conceal a speaker's identity while preserving intelligibility of the underlying speech. We introduce a benchmark that quantifies residual identity leakage with three complementary error rates: equal error rate, cumulative match characteristic hit rate, and embedding-space similarity measured via canonical correlation analysis and Procrustes analysis. Evaluation results reveal that all state-of-the-art speaker de-identification systems leak identity information. The highest performing system in our evaluation performs only slightly better than random guessing, while the lowest performing system achieves a 45% hit rate within the top 50 candidates based on CMC. These findings highlight persistent privacy risks in current speaker de-identification technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。