arXiv:2410.13342eess.AScs.AI2024-10中稿 · NeurIPS被引 9

分离说话人与口音特征,实现任意组合的个性化语音合成

DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech

  • 用多级变分自编码器与向量量化分离说话人和口音表征
  • 在多个数据集上实现更精准的口音模仿与说话人保持
  • 适合需要高度定制化语音输出的应用场景

近年来,文本转语音(TTS)系统已能从文本生成自然且富有表现力的语音。带有口音的TTS旨在提升少数群体听众的使用体验,并广泛适用于各类应用场景。通过允许用户自由组合说话人身份与口音,可进一步提升语音合成的灵活性。然而,现有模型难以有效分离说话人与口音表征,导致在保留原说话人特征的同时准确模仿不同口音存在困难。本文提出一种新方法,利用多级变分自编码器(ML-VAE)与向量量化(VQ)实现说话人与口音表征的解耦,提升语音合成的灵活性与个性化能力。该方法有效分离了说话人与口音特征,实现了更精细的语音控制。代码与语音样本已公开。

原文摘要 · Abstract (English)

Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority group listeners, and useful across various applications and context. Speech synthesis can further be made more flexible by allowing users to choose any combination of speaker identity and accent, resulting in a wide range of personalized speech outputs. Current models struggle to disentangle speaker and accent representation, making it difficult to accurately imitate different accents while maintaining the same speaker characteristics. We propose a novel approach to disentangle speaker and accent representations using multi-level variational autoencoders (ML-VAE) and vector quantization (VQ) to improve flexibility and enhance personalization in speech synthesis. Our proposed method addresses the challenge of effectively separating speaker and accent characteristics, enabling more fine-grained control over the synthesized speech. Code and speech samples are publicly available.

语音合成口音识别表征解耦个性化语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。