arXiv:2506.15614cs.SD2025-06中稿 · IEEE Transactions …被引 6

用闭环优化从杂乱网络语音中自动训练多说话人语音合成模型

TTSOps: A Closed-Loop Corpus Optimization Framework for Training Multi-Speaker TTS Models from Dark Data

  • 基于质量动态选择清洗方法,自适应调整数据处理策略
  • 通过预测的平均意见得分评估语句价值,提升合成自然度和多样性
  • 适合想用海量低质语音数据训练高质量语音模型的研究者

本文提出TTSOps,一个完全自动化、闭环式的框架,用于从嘈杂的网络级语音数据(如在线视频)中构建多说话人文本转语音系统。传统训练依赖高质量、精准对齐的语料,难以扩展且限制多样性。现有方法虽关注音质筛选,却忽视现代语音模型对噪声的鲁棒性及低质但信息丰富的样本价值。TTSOps整合三大核心:(1)从暗数据源自动收集;(2)根据数据质量动态选择逐语句清洗方法;(3)利用自动预测的均值意见分(MOS)进行闭环评估,估算每条语句对模型性能的影响。通过联合优化语料与模型,在日本YouTube数据集上的实验表明,TTSOps在合成语音自然度和说话人多样性上均优于传统音质筛选基线。

原文摘要 · Abstract (English)

This paper presents TTSOps, a fully automated closed-loop framework for constructing multi-speaker text-to-speech (TTS) systems from noisy, uncurated web-scale speech data, often referred to as ``dark data,'' such as online videos. Conventional TTS training pipelines require well-curated corpora with high acoustic quality and accurate text-speech alignment, which severely limits scalability, speaker diversity, and real-world applicability. While recent studies have proposed acoustic-quality-based data selection techniques, they often overlook two critical aspects: (1) the inherent robustness of modern TTS models to noise, and (2) the potential contribution of perceptually low-quality yet informative samples. To address these issues, TTSOps introduces a data-centric training pipeline that integrates three core components: (1) automated data collection from dark data sources, (2) utterance-level dynamic selection of data cleansing methods based on training data quality, and (3) evaluation-in-the-loop data selection using automatically predicted mean opinion scores (MOS) to estimate each utterance's impact on model performance. Furthermore, TTSOps jointly optimizes the corpus and the TTS model in a closed-loop framework by dynamically adapting both data selection and data cleansing processes to the characteristics of the target TTS model. Extensive experiments on Japanese YouTube data demonstrate that TTSOps outperforms conventional acoustic-quality-based baselines in both the naturalness and speaker diversity of synthesized speech.

语音合成数据优化闭环训练多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。