arXiv:2409.03520eess.ASeess.SP2024-09中稿 · EUSIPCO 2024被引 4

分离语音中的说话人与风格特征,无需风格标签即可实现高质量语音转换。

Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder

  • 基于对比预测编码的变分自编码器,利用说话人信息时间稳定性实现初步解耦。
  • 在无风格标签情况下,成功将非内容特征进一步分解为独立的说话人与风格向量。
  • 适用于语音转换任务,尤其适合缺乏标注数据的场景。

语音信号包含内容、说话人和风格等多个层次的信息。尽管解耦这些信息极具挑战性,但在语音转换等应用中至关重要。现有方法基于对比预测编码的因子化变分自编码器,通过假设说话人信息在时间上比内容变化更稳定,实现了语音信号中说话人与内容嵌入的无监督解耦。然而,该假设可能导致其他时间稳定的特征(如环境或情绪)被错误纳入说话人嵌入,这类特征我们称为风格。本文提出一种新方法,进一步将非内容特征解耦为独立的说话人与风格特征,关键在于利用现成且明确的说话人标签,无需依赖风格标签。实验验证了该方法在提取解耦特征方面的有效性,从而支持说话人、风格或联合说话人-风格转换。

原文摘要 · Abstract (English)

Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversion. The contrastive predictive coding supported factorized variational autoencoder achieves unsupervised disentanglement of a speech signal into speaker and content embeddings by assuming speaker info to be temporally more stable than content-induced variations. However, this assumption may introduce other temporal stable information into the speaker embeddings, like environment or emotion, which we call style. In this work, we propose a method to further disentangle non-content features into distinct speaker and style features, notably by leveraging readily accessible and well-defined speaker labels without the necessity for style labels. Experimental results validate the proposed method's effectiveness on extracting disentangled features, thereby facilitating speaker, style, or combined speaker-style conversion.

语音转换特征解耦无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。