arXiv:2409.15882cs.CVeess.SP2024-09被引 3

用VQ-VAE分离语音内容、语调与说话人特征,实现匿名化同时保留情绪。

Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization

  • 三分支结构分别编码内容、语调和说话人特征
  • 合成时结合语调与说话人信息,更好保留情绪
  • 在情绪保持上优于多数基线方法

人类语音包含语调、语言内容和说话人身份。本文提出一种基于向量量化变分自编码器(VQ-VAE)的端到端方法,旨在分离这些语音成分,专门修改说话人身份而保留语言和情感内容。该方法采用三个独立分支分别计算内容、语调和说话人身份的嵌入表示。合成阶段,解码器同时依赖说话人和语调信息进行条件生成,从而捕捉更细微的情感状态并精确调整说话人识别。实验表明,该方法在保持情感信息方面优于大多数基线技术。然而,在其他语音隐私任务上表现有限,表明仍需进一步优化。

原文摘要 · Abstract (English)

Human speech conveys prosody, linguistic content, and speaker identity. This article investigates a novel speaker anonymization approach using an end-to-end network based on a Vector-Quantized Variational Auto-Encoder (VQ-VAE) to deal with these speech components. This approach is designed to disentangle these components to specifically target and modify the speaker identity while preserving the linguistic and emotionalcontent. To do so, three separate branches compute embeddings for content, prosody, and speaker identity respectively. During synthesis, taking these embeddings, the decoder of the proposed architecture is conditioned on both speaker and prosody information, allowing for capturing more nuanced emotional states and precise adjustments to speaker identification. Findings indicate that this method outperforms most baseline techniques in preserving emotional information. However, it exhibits more limited performance on other voice privacy tasks, emphasizing the need for further improvements.

语音匿名化VQ-VAE情感保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。