提出双流攻击模型,有效破解语音匿名化隐私保护。
DAST: A Dual-Stream Voice Anonymization Attacker with Staged Training
- 双流结构融合频谱与自监督特征,分三阶段训练增强攻击能力。
- 第二阶段使模型在未见数据上表现更优,是泛化关键。
- 仅用10%目标数据微调,超越当前最优攻击方法。
语音匿名化在保留语义内容的同时隐藏声音特征,但仍可能泄露说话人特定模式。为评估并强化隐私安全性,本文提出一种双流攻击模型,通过并行编码器融合频谱特征与自监督学习特征,并采用三阶段训练策略:第一阶段建立基础的说话人判别表征;第二阶段利用语音转换与匿名化的共性,引入多样化转换语音以提升跨系统鲁棒性;第三阶段对目标匿名数据进行轻量级适配。在VoicePrivacy Attacker Challenge(VPAC)数据集上的结果表明,第二阶段是泛化性能的主要驱动力,使模型在未见过的匿名化数据集上仍具备强攻击能力。结合第三阶段,仅用10%的目标数据微调,即可在EER指标上超越现有最先进攻击模型。
原文摘要 · Abstract (English)
Voice anonymization masks vocal traits while preserving linguistic content, which may still leak speaker-specific patterns. To assess and strengthen privacy evaluation, we propose a dual-stream attacker that fuses spectral and self-supervised learning features via parallel encoders with a three-stage training strategy. Stage I establishes foundational speaker-discriminative representations. Stage II leverages the shared identity-transformation characteristics of voice conversion and anonymization, exposing the model to diverse converted speech to build cross-system robustness. Stage III provides lightweight adaptation to target anonymized data. Results on the VoicePrivacy Attacker Challenge (VPAC) dataset demonstrate that Stage II is the primary driver of generalization, enabling strong attacking performance on unseen anonymization datasets. With Stage III, fine-tuning on only 10\% of the target anonymization dataset surpasses current state-of-the-art attackers in terms of EER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。