构建22种印度语言的统一语音合成数据集,小规模高质量数据即可训练出优秀模型
A Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages
- 建立覆盖22种印度语言的标准化数据采集框架
- 少量纯净数据即能支持高保真语音合成,优于欧美/中文数据集
- 适用于多语言语音合成系统开发与学术研究
语音合成模型性能受多种因素影响,其中训练数据质量至关重要。全球已积累海量数据资源,但印度语言数据仍极为稀缺。尽管已有诸多数据收集努力,但为实现各语言间一致的语音合成系统开发,建立统一的数据采集协议势在必行。本文总结了过去15年针对印地语系语言的数据收集经验,这些数据已被用于单元选择合成、隐马尔可夫模型及端到端框架,并支持生成富有韵律感的语音合成系统。所收集数据最显著的特点是数据纯净度高,仅需相对较少的数据量即可构建高质量语音合成系统,相较于欧洲或中文语言数据集更具效率。
原文摘要 · Abstract (English)
The performance of a text-to-speech (TTS) synthesis model depends on various factors, of which the quality of the training data is of utmost importance. Millions of data are collected around the globe for various languages, but resources for Indian languages are few. Although there are many efforts involved in data collection, a common set of protocols for data collection becomes necessary for building TTS systems in Indian languages primarily because of the need for a uniform development of TTS systems across languages. In this paper, we present our learnings on data collection efforts' for Indic languages over 15 years. These databases have been used in unit selection synthesis, hidden Markov model based, and end-to-end frameworks, and for generating prosodically rich TTS systems. The most significant feature of the data collected is that data purity enables building high-quality TTS systems with a comparatively small dataset compared to that of European/Chinese languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。