arXiv:2602.10164cs.SDeess.AS2026-02中稿 · IEEE Spoken Langua…

用情感一致语音增强小数据,提升儿童故事语音合成自然度

Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis

  • 通过情感匹配文本合并音频,生成更丰富的表达性语音数据
  • 模型生成的句子间停顿分布更接近真实语音,主观评分更高
  • 适合需要高质量语音合成的儿童内容生成场景

富有表现力的语音合成需要生动的语调和恰当的停顿。本文提出一种有效策略,通过情感识别器筛选情感一致的文本,合并对应音频以扩充小规模数据集,训练端到端语音合成模型。利用双句音频进行训练,使模型学习到句子间的自然停顿。进一步引入自监督对比学习,优化语音中说话风格嵌入的提取。推理时,模型可一次性生成多句语音,由文本预测的风格引导。实验表明,相比仅使用连续双句音频训练的基线模型,本文方法在句子间停顿分布上更接近真实语音,主观评测显示其自然度和风格适配性均显著提升。

原文摘要 · Abstract (English)

Expressive speech synthesis requires vibrant prosody and well-timed pauses. We propose an effective strategy to augment a small dataset to train an expressive end-to-end Text-to-Speech model. We merge audios of emotionally congruent text using a text emotion recognizer, creating augmented expressive speech data. By training with two-sentence audio, our model learns natural breaks between lines. We further apply self-supervised contrastive training to improve the speaking style embedding extraction from speech. During inference, our model produces multi-sentence speech in one step, guided by the text-predicted speaking style. Evaluations showcase the effectiveness of our proposed approach when compared to a baseline model trained with consecutive two-sentence audio. Our synthesized speeches give a closer inter-sentence pause distribution to the ground truth speech. Subjective evaluations reveal our synthesized speech scored higher in naturalness and style suitability than the baseline.

语音合成情感增强数据增强自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。