无需外部编码器,实现自然且稳定的跨说话人情感风格迁移。
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
- 用梯度反向层和余弦相似性损失分离说话人与情感特征。
- 通过多正例对比学习使说话人和情感嵌入聚类分明。
- 自增强策略提升语音自然度,适合语音合成与风格迁移研究者。
本文提出SelfTTS,一种用于跨说话人风格迁移的文本到语音(TTS)模型,无需依赖外部预训练的说话人或情感编码器。该架构通过显式解耦策略,利用梯度反向层(GRL)结合余弦相似性损失,分离说话人与情感信息,使中性语调的说话人具备情感表现力。我们引入多正例对比学习(MPCL),基于标签引导说话人与情感嵌入形成聚类表示。此外,SelfTTS采用自增强策略,利用模型自身语音转换能力进行自我优化,提升合成语音的自然度。实验表明,SelfTTS在情感自然度(eMOS)和目标音色、情感稳定性方面均优于现有最先进基线方法。
原文摘要 · Abstract (English)
This paper presents SelfTTS, a text-to-speech (TTS) model designed for cross-speaker style transfer that eliminates the need for external pre-trained speaker or emotion encoders. The architecture achieves emotional expressivity in neutral speakers through an explicit disentanglement strategy utilizing Gradient Reversal Layers (GRL) combined with cosine similarity loss to decouple speaker and emotion information. We introduce Multi Positive Contrastive Learning (MPCL) to induce clustered representations of speaker and emotion embeddings based on their respective labels. Furthermore, SelfTTS employs a self-refinement strategy via Self-Augmentation, exploiting the model's voice conversion capabilities to enhance the naturalness of synthesized speech. Experimental results demonstrate that SelfTTS achieves superior emotional naturalness (eMOS) and robust stability in target timbre and emotion compared to state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。