arXiv:2608.15910eess.AScs.CL2026-08

用迭代自学习解决语音合成中情感标签稀缺问题。

Iterative Self-Learning for Expressive Text-to-Speech Synthesis

论文配图:Iterative Self-Learning for Expressive Text-to-Speech Synthesis
图 1 · 摘自论文原文
  • 通过反向生成模型自动为无标签语音打情感标签。
  • 多轮迭代提升伪标签准确率,使合成语音更贴合情感要求。
  • 适合数据少但需高表达力语音合成的场景。

强调性文本到语音(TTS)系统若使用显式条件标签,可直接控制语音表现力,但依赖标注数据。大规模获取此类标签成本高昂,而现有半监督方法未解决此瓶颈。本文提出一种基于反向分类(Invert-Classify)的迭代自学习(ISL)框架,通过冻结生成模型反推离散表达标签,对无标签语音进行伪标注。模型在有标签与伪标签数据上反复训练,逐步优化标签质量与合成效果。在词级重音和句级情感两个任务上验证,结果表明迭代优化显著提升伪标签精度,并带来更好的情感遵循度与合成质量,客观指标与人工听感均证实其有效性。在数据最稀缺情况下,该方法性能接近全监督模型,证明梯度驱动的迭代自学习是应对低资源表达性语音合成中标签稀缺的有效方案。

原文摘要 · Abstract (English)

Expressive text-to-speech (TTS) systems that use explicit conditioning labels provide direct and interpretable control over expressive attributes, in contrast to reference-based or prompting-based approaches, but require labeled data. Obtaining these labels at scale is costly and time-consuming, yet no prior semi-supervised framework addresses this specific bottleneck. Existing semi-supervised TTS methods instead target scarcity of paired speech-text data or transcriptions. To address the scarcity of expressive labels, we propose an Iterative Self-Learning (ISL) framework for expressive TTS, built on Invert-Classify, a classifier-free method that recovers discrete expressive labels by inverting a frozen generative model. The framework iteratively pseudo-labels unlabeled speech using the current model, retrains on the combined labeled and pseudo-labeled data, and repeats, progressively refining label quality and synthesis. We validate on two expressive tasks, word-level prominence and utterance-level emotion, across multiple low-resource data splits. We find that iterative refinement can improve pseudo-label accuracy over single-pass baselines. Furthermore, we observe that these improvements in pseudo-labeling of expressivity translate to gains in expressive label adherence and synthesis quality, confirmed by objective metrics and human listening tests. In the most data-scarce conditions, ISL-trained models outperform single-pass pseudo-labeling and further approach fully supervised performance, demonstrating that gradient-based ISL is an effective solution to expressive label scarcity in low-resource TTS.

语音合成自学习低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。