用索引映射生成唯一语音身份,提升隐私保护效率。
IDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector Mapping
- 通过索引直接映射到声纹向量,无需复杂生成过程。
- 在小规模数据上提升伪说话人唯一性,计算成本更低。
- 大规模场景下仍稳定保持隐私保护效果,适合实际应用。
基于内容、说话人和语调解耦的语音生成框架,可通过替换原始说话人嵌入向量实现语音匿名化。其中,伪说话人生成是核心挑战。现有方法在伪说话人唯一性方面存在不足,且模型方法计算开销大。尤其在需生成大量伪说话人的场景下,唯一性和效率问题更为突出。为此,本文提出一种基于说话人身份索引到向量映射的前馈式框架IDMap,包含IDMap-MLP与IDMap-Diff两种模型。在小规模数据集LibriSpeech上的实验验证了该框架在提升伪说话人唯一性的同时降低计算成本,增强了语音隐私保护能力;在大规模数据集MLS和Common Voice上的实验进一步证明,随着伪说话人数量增加,其隐私保护能力保持稳定。音频样本与开源代码见https://github.com/VoicePrivacy/IDMap。
原文摘要 · Abstract (English)
Facilitated by the speech generation framework that disentangles speech into content, speaker, and prosody, voice anonymization is accomplished by substituting the original speaker embedding vector with that of a pseudo-speaker. In this framework, the pseudo-speaker generation forms a fundamental challenge. Current pseudo-speaker generation methods demonstrate limitations in the uniqueness of pseudo-speakers, consequently restricting their effectiveness in voice privacy protection. Besides, existing model-based methods suffer from heavy computation costs. Especially, in the large-scale scenario where a huge number of pseudo-speakers are generated, the limitations of uniqueness and computational inefficiency become more significant. To this end, this paper proposes a framework for pseudo-speaker generation, which establishes a mapping from speaker identity index to speaker vector in the feedforward architecture, termed IDMap. Specifically, the framework is specified into two models: IDMap-MLP and IDMap-Diff. Experiments were conducted on both small- and large-scale evaluation datasets. Small-scale evaluations on the LibriSpeech dataset validated the effectiveness of the proposed IDMap framework in enhancing the uniqueness of pseudo-speakers, thereby improving voice privacy protection, while at a reduced computational cost. Large-scale evaluations on the MLS and Common Voice datasets further justified the superiority of the IDMap framework regarding the stability of the voice privacy protection capability as the number of pseudo-speakers increased. Audio samples and open-source code can be found in https://github.com/VoicePrivacy/IDMap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。