arXiv:2509.00675eess.AS2025-09

用音素级预训练模型提升多说话人语音合成的语义断句准确率

Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model

  • 引入说话人嵌入,让模型学习不同说话人的语义断句特征
  • 音素级预训练模型使断句准确率显著提升,超越传统方法
  • 支持对未见说话人进行少量样本适应,适合实际语音合成场景

本文推进多说话人文本到语音系统中的语义断句(即分句)任务。通过融合说话人嵌入,利用说话人特定特征提升分句模型性能,并证明这些嵌入可仅从分句任务中捕捉说话人相关特征。此外,我们探索了预训练说话人嵌入在未见说话人上的少样本适应能力。更重要的是,我们首次将音素级预训练语言模型应用于该前端任务,显著提升了分句模型的准确性。我们的方法通过客观与主观评估进行了严格验证,结果表明其有效性。

原文摘要 · Abstract (English)

This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We further demonstrate that these speaker embeddings can capture speaker-related characteristics solely from the phrasing task. Besides, we explore the potential of pre-trained speaker embeddings for unseen speakers through a few-shot adaptation method. Furthermore, we pioneer the application of phoneme-level pre-trained language models to this TTS front-end task, which significantly boosts the accuracy of the phrasing model. Our methods are rigorously assessed through both objective and subjective evaluations, demonstrating their effectiveness.

语音合成语义断句说话人嵌入预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。