构建超大规模多语言口语数据集,提升语音生成自然度。
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
- 从真实场景音频中提取高质量语音数据,突破传统朗读式语料局限。
- 建成101小时跨六语言语音数据集,扩展后达216小时以上。
- 适合多语言、跨语言语音生成研究,推动更自然的对话系统发展。
近期语音生成的进步依赖于大规模训练数据集。然而,现有模型难以捕捉真实人类语音中的自发性和多样性,因其主要在以正式朗读为主的有声书数据集上训练。为解决这一问题,我们提出Emilia-Pipe——一个开源预处理管道,用于从潜在高价值但未被充分挖掘的真实场景音频中提取高质量训练数据。利用该工具,我们构建了包含六种语言(英语、中文、德语、法语、日语、韩语)超过101,000小时语音的Emilia数据集。进一步扩展至超过216,000小时的Emilia-Large,成为目前最大规模的开源语音生成资源之一。大量实验表明,基于Emilia训练的模型生成语音更具自发性与人类感,同时保持与传统数据集相当的可懂度;其更能体现多样化的说话人音色和真实对话风格。本工作验证了数据规模扩增对语音生成性能提升的重要性,并证明Emilia在多语言及跨语言语音生成任务中的有效性。
原文摘要 · Abstract (English)
Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in real-world human speech, as they are primarily trained on audio-book datasets limited to formal, read-aloud speaking styles. To address this limitation, we introduce Emilia-Pipe, an open-source preprocessing pipeline designed to extract high-quality training data from valuable yet under-explored in-the-wild sources that capture spontaneous human speech in real-world contexts. Using Emilia-Pipe, we construct Emilia, which comprises over 101k hours of speech across six languages: English, Chinese, German, French, Japanese, and Korean. Furthermore, we expand Emilia to Emilia-Large, a dataset exceeding 216k hours, making it one of the largest open-source speech generation resources available. Extensive experiments show that Emilia-trained models produce markedly more spontaneous, human-like speech than those trained on traditional audio-book datasets, while matching their intelligibility. These models better capture diverse speaker timbres and the full spectrum of real-world conversational styles. Our work also highlights the importance of scaling dataset size for advancing speech generation performance and validates the effectiveness of Emilia for both multilingual and crosslingual speech generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。