开源芬兰语和瑞典语语音数据集,助力低资源语言语音合成
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
- 基于北欧议会录音构建,自动化处理生成高质量语音数据
- 含900小时芬兰语和5090小时瑞典语语音,支持模型训练与评估
- 为低资源语言语音合成提供大规模公开数据,适合语音研究者使用
文本到语音(TTS)发展受限于多数语言缺乏高质量、公开可用的语音数据,尤其是高资源语言之外的语言。本文提出 Nord-Parl-TTS,一个基于真实场景录音的芬兰语和瑞典语开放 TTS 数据集。利用北欧议会会议录音,我们提取了900小时芬兰语和5090小时瑞典语语音,适用于 TTS 训练。该数据集采用改进版 Emilia 数据处理流程构建,并包含统一的评估集,支持模型开发与基准测试。通过提供大规模、公开的芬兰语和瑞典语数据,Nord-Parl-TTS 缩小了高资源与低资源语言在语音合成领域的数据差距。
原文摘要 · Abstract (English)
Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages. We present Nord-Parl-TTS, an open TTS dataset for Finnish and Swedish based on speech found in the wild. Using recordings of Nordic parliamentary proceedings, we extract 900 hours of Finnish and 5090 hours of Swedish speech suitable for TTS training. The dataset is built using an adapted version of the Emilia data processing pipeline and includes unified evaluation sets to support model development and benchmarking. By offering open, large-scale data for Finnish and Swedish, Nord-Parl-TTS narrows the resource gap in TTS between high- and lower-resourced languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。