通过批量伪说话人实现任意说话人匿名,保护隐私不依赖真实目标说话人。
Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization
- 用批内平均说话人作为伪目标,实现任意说话人间的属性扰动。
- 在VoxCeleb数据集上验证了匿名化效果,有效提升身份不可关联性。
- 适合关注语音隐私保护、对抗黑盒提取攻击的研究者使用。
说话人属性扰动为异步语音匿名化提供可行方案,通过对抗性扰动语音生成匿名输出。为增强同一原始说话人生成的匿名语句之间的身份不可关联性,通常采用目标攻击训练策略,将语句匿名化至指定目标说话人。然而该策略可能侵犯目标说话人的隐私。为此,本文提出任意到任意训练策略:定义批量均值损失,将训练小批次中不同说话人的语音匿名化至一个由批次内平均说话人近似得到的伪说话人。基于此,提出一种说话人对抗性语音生成模型,同时结合无目标攻击和任意到任意策略的监督。生成的说话人属性扰动被注入原始语音,生成匿名版本。实验在VoxCeleb数据集上验证了所提模型在异步语音匿名化中的有效性。额外实验探讨了说话人对抗性语音在语音隐私保护中的潜在局限性,旨在为未来研究提供关于其对抗黑盒说话人提取器与自适应攻击的防护能力、跨域泛化性及稳定性等方面的洞见。音频样本与开源代码已发布于 https://github.com/VoicePrivacy/any-to-any-speaker-attribute-perturbation。
原文摘要 · Abstract (English)
Speaker attribute perturbation offers a feasible approach to asynchronous voice anonymization by employing adversarially perturbed speech as anonymized output. In order to enhance the identity unlinkability among anonymized utterances from the same original speaker, the targeted attack training strategy is usually applied to anonymize the utterances to a common designated speaker. However, this strategy may violate the privacy of the designated speaker who is an actual speaker. To mitigate this risk, this paper proposes an any-to-any training strategy. It is accomplished by defining a batch mean loss to anonymize the utterances from various speakers within a training mini-batch to a common pseudo-speaker, which is approximated as the average speaker in the mini-batch. Based on this, a speaker-adversarial speech generation model is proposed, incorporating the supervision from both the untargeted attack and the any-to-any strategies. The speaker attribute perturbations are generated and incorporated into the original speech to produce its anonymized version. The effectiveness of the proposed model was justified in asynchronous voice anonymization through experiments conducted on the VoxCeleb datasets. Additional experiments were carried out to explore the potential limitations of speaker-adversarial speech in voice privacy protection. With them, we aim to provide insights for future research on its protective efficacy against black-box speaker extractors \textcolor{black}{and adaptive attacks, as well as} generalization to out-of-domain datasets \textcolor{black}{and stability}. Audio samples and open-source code are published in https://github.com/VoicePrivacy/any-to-any-speaker-attribute-perturbation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。