arXiv:2509.14270cs.CLcs.AI2025-09ACL被引 7

SpeechWeave自动化生成多语言语音数据,提升多样性与文本规范性。

SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models

  • 用合成方法自动构建多语言、领域特定的文本与语音数据集。
  • 生成数据在语言和语音特征上比基线多样10%-48%,文本正确归一化率达97%。
  • 适合需要高质量、标准化语音训练数据的TTS系统研发团队。

高质量文本转语音(TTS)模型训练需要大量且多样化的文本与语音数据。然而,真实数据受限于领域特异性、版权问题及可扩展性,难以获取。大语言模型虽能生成文本,但常导致重复内容且提示变化不足。文本归一化工具也可能引入异常或遗漏有效模式,影响数据质量。此外,商业TTS系统依赖配音艺术家大规模录音不现实。为此,我们提出SpeechWeave,一个可自动化生成多语言、领域特定语音数据的合成流水线。实验表明,该流水线生成的数据在多种语言与音素度量下多样性提升10%-48%,同时生成约97%正确归一化的文本,并实现说话人标准化语音输出。该方法显著提升了TTS训练数据的多样性、归一化准确率与语音一致性。

原文摘要 · Abstract (English)

High-quality Text-to-Speech (TTS) model training requires extensive and diverse text and speech data. It is challenging to procure such data from real sources due to issues of domain specificity, licensing, and scalability. Large language models (LLMs) can certainly generate textual data, but they create repetitive text with insufficient variation in the prompt during the generation process. Another important aspect in TTS training data is text normalization. Tools for normalization might occasionally introduce anomalies or overlook valuable patterns, and thus impact data quality. Furthermore, it is also impractical to rely on voice artists for large scale speech recording in commercial TTS systems with standardized voices. To address these challenges, we propose SpeechWeave, a synthetic speech data generation pipeline that is capable of automating the generation of multilingual, domain-specific datasets for training TTS models. Our experiments reveal that our pipeline generates data that is 10-48% more diverse than the baseline across various linguistic and phonetic metrics, along with speaker-standardized speech audio while generating approximately 97% correctly normalized text. Our approach enables scalable, high-quality data generation for TTS training, improving diversity, normalization, and voice consistency in the generated datasets.

语音合成数据生成多语言文本归一化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。