arXiv:2507.08319cs.SDeess.AS2025-07被引 1

用主动学习选更有价值的语音数据,省空间还更清晰。

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection

  • 边训练边选数据,只留对模型提升最有帮助的样本
  • 同等规模下,合成语音质量显著优于传统方法
  • 适合资源有限但追求高质量语音的开发者

高质量数据集是现代文语转换(TTS)系统的基础。然而,可用数据规模不断增长,带来了存储压力。为应对这一挑战,我们提出一种基于主动学习的TTS语料构建方法。与传统的前馈式、模型无关的语料构建方法不同,该方法在数据收集与模型训练之间迭代交替,聚焦于获取对模型改进更具信息量的数据。此方法可实现数据高效的语料构建。实验结果表明,使用该方法构建的语料库,在相同规模下能实现比传统语料库更高质量的语音合成。

原文摘要 · Abstract (English)

The construction of high-quality datasets is a cornerstone of modern text-to-speech (TTS) systems. However, the increasing scale of available data poses significant challenges, including storage constraints. To address these issues, we propose a TTS corpus construction method based on active learning. Unlike traditional feed-forward and model-agnostic corpus construction approaches, our method iteratively alternates between data collection and model training, thereby focusing on acquiring data that is more informative for model improvement. This approach enables the construction of a data-efficient corpus. Experimental results demonstrate that the corpus constructed using our method enables higher-quality speech synthesis than corpora of the same size.

TTS主动学习数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。