通过自监督蒸馏分离情感与音色,实现跨说话人情感迁移的高质量语音合成。
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
- 采用自监督蒸馏与聚类采样,分离情感与说话人特征。
- 在未标注数据上实现情感无关嵌入,减少音色泄露。
- 适合需要高保真情感迁移的语音合成场景。
语音合成中的跨说话人情感迁移依赖于提取与说话人无关的情感嵌入以实现精准情感建模,同时不保留说话人特征。然而,现有音色压缩方法无法完全分离说话人与情感特征,导致音色泄露和合成质量下降。为此,我们提出DiEmo-TTS,一种自监督蒸馏方法,旨在最小化情感信息损失并保留说话人身份。引入基于聚类的采样与信息扰动机制,在保留情感特征的同时去除无关因素。为支持该过程,我们设计了一种结合情感属性预测与说话人嵌入的情感聚类匹配方法,可推广至未标注数据。此外,我们还构建了双条件变压器以更优整合风格特征。实验结果验证了该方法在学习与说话人无关的情感嵌入方面的有效性。
原文摘要 · Abstract (English)
Cross-speaker emotion transfer in speech synthesis relies on extracting speaker-independent emotion embeddings for accurate emotion modeling without retaining speaker traits. However, existing timbre compression methods fail to fully separate speaker and emotion characteristics, causing speaker leakage and degraded synthesis quality. To address this, we propose DiEmo-TTS, a self-supervised distillation method to minimize emotional information loss and preserve speaker identity. We introduce cluster-driven sampling and information perturbation to preserve emotion while removing irrelevant factors. To facilitate this process, we propose an emotion clustering and matching approach using emotional attribute prediction and speaker embeddings, enabling generalization to unlabeled data. Additionally, we designed a dual conditioning transformer to integrate style features better. Experimental results confirm the effectiveness of our method in learning speaker-irrelevant emotion embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。