arXiv:2606.15267eess.AScs.SD2026-06中稿 · INTERSPEECH 2026

通过动态预测语调提升大模型语音合成的发音风格相似性

Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity

  • 根据已生成语音动态预测当前音节语调,强化风格学习
  • 在三个数据集上显著提升语音合成的说话人相似度
  • 适合追求高保真语音克隆的个性化合成研究者

个性化文本转语音(TTS)旨在克隆目标说话人的语音与表达风格。现有基于大语言模型(LLM)的TTS方法忽略生成语音中的风格特异性语调模式,导致风格学习不足,限制了合成语音的说话人相似度。为此,本文研究了基于生成语音的语调学习机制,提出根据先前预测的语音动态预测当前音节的语调。在三个数据集上的实验结果表明,所提出的动态语调预测方法有效增强了语调学习能力,从而提升了合成语音的说话人相似度。音频样例见 https://muzw.github.io/dynapros/。

原文摘要 · Abstract (English)

Personalized text-to-speech (TTS) aims to clone the target speaker in the synthesized speech, imitating both the voice and speaking style. Current large language model (LLM)-based TTS methods ignore the style-specific prosodic patterns in generated speech, resulting in deficient style learning and thus limiting speaker similarity in synthesized speech. To this end, we investigate the prosody learning conditioned on the synthesized speech, and propose to predict the prosody of the current syllable based on previously predicted speech. Experimental results obtained on three datasets demonstrated the efficacy of the proposed dynamic prosody prediction method in enhancing the prosody learning capability, thereby improving the speaker similarity of the generated speech. Audio samples are available at https://muzw.github.io/dynapros/.

语音合成大模型风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。