DurIAN-E 2 用自适应变分自编码与对抗训练,提升语音合成的自然度和表现力。
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
- 结合变分自编码器与归一化流,增强语音表征学习能力。
- 采用大尺度对抗训练的 BigVGAN 生成器,合成语音质量更高。
- 适合追求高保真、情感丰富语音合成的研究者与开发者。
本文提出 DurIAN-E 2,一种用于表达性高质量文语转换(TTS)的改进型时长感知注意力网络。与原始 DurIAN-E 类似,该模型采用多个堆叠的 SwishRNN 基 Transformer 模块作为语言编码器,并在帧级编码器中引入风格自适应实例归一化(SAIN)层,以增强表现力建模能力。同时,受 VITS 等生成式模型启发,DurIAN-E 2 引入带归一化流的变分自编码器(VAEs)与采用对抗训练策略的大规模 BigVGAN 波形生成器,进一步提升合成语音的质量与表现力。客观测试与主观评估结果表明,该模型在多项指标上优于多个先进方法,包括原始 DurIAN-E。
原文摘要 · Abstract (English)
This paper proposes an improved version of DurIAN-E (DurIAN-E 2), which is also a duration informed attention neural network for expressive and high-fidelity text-to-speech (TTS) synthesis. Similar with the DurIAN-E model, multiple stacked SwishRNN-based Transformer blocks are utilized as linguistic encoders and Style-Adaptive Instance Normalization (SAIN) layers are also exploited into frame-level encoders to improve the modeling ability of expressiveness in the proposed the DurIAN-E 2. Meanwhile, motivated by other TTS models using generative models such as VITS, the proposed DurIAN-E 2 utilizes variational autoencoders (VAEs) augmented with normalizing flows and a BigVGAN waveform generator with adversarial training strategy, which further improve the synthesized speech quality and expressiveness. Both objective test and subjective evaluation results prove that the proposed expressive TTS model DurIAN-E 2 can achieve better performance than several state-of-the-art approaches besides DurIAN-E.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。