15万小时多语言语音数据集,支持零样本合成与自然编辑
LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models
- 用生成模型构建15万小时跨10语种高精度标注数据集
- 零样本多语言合成与无缝语音编辑均达高质量表现
- 适合语音生成、语音编辑及多语言模型研究者使用
我们提出LEMAS-Dataset,据我们所知,目前最大开源的多语言语音语料库,包含超过15万小时数据,覆盖10种主要语言,具有词级时间戳。该数据集通过高效数据处理流程构建,确保高质量与精准标注。为验证其在多种生成范式下的有效性,我们在该数据集上训练了两个不同架构的基准模型:LEMAS-TTS基于非自回归流匹配框架,利用大规模多样语料实现稳健的零样本多语言语音合成;通过提出口音对抗训练与CTC损失,缓解跨语言口音问题,提升合成稳定性。同时,LEMAS-Edit采用自回归解码器架构,将语音编辑建模为掩码标记填充任务,借助精确词级对齐构造训练掩码并采用自适应解码策略,实现自然过渡的平滑语音编辑。实验表明,基于LEMAS-Dataset训练的模型在合成与编辑任务中均表现优异,验证了数据集质量。我们期望这一细粒度、带时间戳的多语言语料库能推动提示驱动语音生成系统的发展。
原文摘要 · Abstract (English)
We present the LEMAS-Dataset, which, to our knowledge, is currently the largest open-source multilingual speech corpus with word-level timestamps. Covering over 150,000 hours across 10 major languages, LEMAS-Dataset is constructed via a efficient data processing pipeline that ensures high-quality data and annotations. To validate the effectiveness of LEMAS-Dataset across diverse generative paradigms, we train two benchmark models with distinct architectures and task specializations on this dataset. LEMAS-TTS, built upon a non-autoregressive flow-matching framework, leverages the dataset's massive scale and linguistic diversity to achieve robust zero-shot multilingual synthesis. Our proposed accent-adversarial training and CTC loss mitigate cross-lingual accent issues, enhancing synthesis stability. Complementarily, LEMAS-Edit employs an autoregressive decoder-only architecture that formulates speech editing as a masked token infilling task. By exploiting precise word-level alignments to construct training masks and adopting adaptive decoding strategies, it achieves seamless, smooth-boundary speech editing with natural transitions. Experimental results demonstrate that models trained on LEMAS-Dataset deliver high-quality synthesis and editing performance, confirming the dataset's quality. We envision that this richly timestamp-annotated, fine-grained multilingual corpus will drive future advances in prompt-based speech generation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。