arXiv:2507.12081eess.AS2025-07中稿 · WASPAA 2025被引 1

用语音和文字联合攻击语音匿名系统,暴露其隐私漏洞。

VoxATtack: A Multimodal Attack on Voice Anonymization Systems

  • 结合语音与文本信息构建双分支攻击模型
  • 在VPAC数据集上5项指标超越顶尖攻击者
  • 揭示现有匿名技术与评测数据集的潜在缺陷

语音匿名系统通过隐藏声纹特征来保护说话人隐私,同时保留语义内容以支持下游应用。然而,因语义线索仍可被利用,攻击者可通过特定说话人关联的语义模式进行识别。本文提出VoxATtack,一种融合声学与文本信息的多模态去匿名化模型。该模型采用双分支结构:ECAPA-TDNN处理匿名化语音,预训练BERT编码转录文本。两者输出映射至同维嵌入,并基于每句话的置信度加权融合。在VoicePrivacy Attacker Challenge(VPAC)数据集上,其在五项基准测试(B3、B4、B5、T8-5、T12-5)中表现优于当前最佳攻击者。引入匿名语音与SpecAugment增强后,进一步提升性能,在T10-2与T25-1上分别达到20.6%与27.2%的平均等错误率,实现全基准领先。结果表明,结合文本信息与选择性数据增强能有效揭示当前语音匿名方法的关键弱点,并暴露评测数据集的潜在缺陷。

原文摘要 · Abstract (English)

Voice anonymization systems aim to protect speaker privacy by obscuring vocal traits while preserving the linguistic content relevant for downstream applications. However, because these linguistic cues remain intact, they can be exploited to identify semantic speech patterns associated with specific speakers. In this work, we present VoxATtack, a novel multimodal de-anonymization model that incorporates both acoustic and textual information to attack anonymization systems. While previous research has focused on refining speaker representations extracted from speech, we show that incorporating textual information with a standard ECAPA-TDNN improves the attacker's performance. Our proposed VoxATtack model employs a dual-branch architecture, with an ECAPA-TDNN processing anonymized speech and a pretrained BERT encoding the transcriptions. Both outputs are projected into embeddings of equal dimensionality and then fused based on confidence weights computed on a per-utterance basis. When evaluating our approach on the VoicePrivacy Attacker Challenge (VPAC) dataset, it outperforms the top-ranking attackers on five out of seven benchmarks, namely B3, B4, B5, T8-5, and T12-5. To further boost performance, we leverage anonymized speech and SpecAugment as augmentation techniques. This enhancement enables VoxATtack to achieve state-of-the-art on all VPAC benchmarks, after scoring 20.6% and 27.2% average equal error rate on T10-2 and T25-1, respectively. Our results demonstrate that incorporating textual information and selective data augmentation reveals critical vulnerabilities in current voice anonymization methods and exposes potential weaknesses in the datasets used to evaluate them.

语音匿名多模态攻击隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。