让语音合成同时保留说话人特征并精准模仿多种口音。
Joycent: Multi-Accent TTS via Disentangled Accent Modeling and Layer-Specific Conditioning

- 分离口音与说话人特征,分层注入生成过程。
- 在跨口音场景下仍保持高口音相似度与说话人一致性。
- 适合需要多口音语音合成的研究者与开发者。
口音文本转语音(TTS)旨在合成具有目标口音的同时保留说话人身份,但面临两大挑战:分离口音与说话人特征,以及有效利用这两个解耦因素进行语音生成。本文提出基于扩散模型的Joycent框架,通过在表征学习和条件生成中分离口音与说话人信息来应对上述问题。采用基于Whisper的口音编码器WhisAID,并使用梯度反转学习去相关化的口音表示;引入分层条件层归一化,在文本编码器不同层级注入口音与说话人信息。在包含已见与未见说话人的普通话地域口音语料库(MRAC)上评估,包括说话人与口音提示来自不同口音的挑战性跨口音设置。实验结果表明,相较于现有方法,Joycent在保持强说话人相似度的同时显著提升口音相似度,且在跨口音设置下持续获得优势。主观评价进一步验证了其在自然度、口音相似度与说话人保留方面的改进。音频样例可访问 https://oshindow.github.io/joycent/。
原文摘要 · Abstract (English)
Accent text-to-speech (TTS) aims to synthesize speech with a target accent while preserving speaker identity, but faces two key challenges: disentangling accent from speaker characteristics and effectively conditioning speech generation on the two disentangled factors. In this paper, we propose Joycent, a diffusion-based accent TTS framework that addresses both challenges. Our key idea is to separate accent and speaker information in both representation learning and TTS conditioning. Joycent uses WhisAID, a Whisper-based accent encoder with gradient reversal to learn speaker-disentangled accent representations, and introduces layer-specific conditional layer normalization to inject accent and speaker information at different stages of the text encoder. We evaluate Joycent on the Mandarin Regional Accent Corpus (MRAC) with seen and unseen speakers, including a challenging cross-accent setting where the speaker and accent prompts come from different accents. Experimental results show that Joycent improves accent similarity over existing methods while maintaining strong speaker similarity, with consistent gains under the challenging cross-accent setting. Subjective evaluation further confirms improved naturalness, accent similarity, and speaker preservation. The audio samples are available at https://oshindow.github.io/joycent/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。