arXiv:2410.06608cs.SDcs.AI2024-10EMNLP被引 1

构建首个大规模印尼语语音合成数据集,提升语音自然度与实用性。

Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS

  • 采用离散编码器建模,提升印尼语语音合成质量
  • 达成4.45分平均主观评分,优于现有基线模型
  • 适合多场景应用,兼顾实时性与模型效率

本研究提出一个全面的印尼语文本转语音(TTS)数据集及新型模型EnGen-TTS,旨在提升印尼语合成语音的质量与多样性。数据集涵盖约55.0小时、52,000段音频,整合多种文本来源,确保语言丰富性;通过专业设备进行高保真录音,捕捉印尼语发音细节。统计分析表明该数据集规模大且多样化,为模型训练与评估奠定基础。所提出的EnGen-TTS模型在多项指标上优于现有基线,实现4.45±0.13的平均意见分数(MOS)。同时,对实时性与模型大小的分析显示其具备高效表现,是极具潜力的实用方案。相关生成样例可访问:https://bahasa-harmony-comp.vercel.app/

原文摘要 · Abstract (English)

This research introduces a comprehensive Bahasa text-to-speech (TTS) dataset and a novel TTS model, EnGen-TTS, designed to enhance the quality and versatility of synthetic speech in the Bahasa language. The dataset, spanning \textasciitilde55.0 hours and 52K audio recordings, integrates diverse textual sources, ensuring linguistic richness. A meticulous recording setup captures the nuances of Bahasa phonetics, employing professional equipment to ensure high-fidelity audio samples. Statistical analysis reveals the dataset's scale and diversity, laying the foundation for model training and evaluation. The proposed EnGen-TTS model performs better than established baselines, achieving a Mean Opinion Score (MOS) of 4.45 $\pm$ 0.13. Additionally, our investigation on real-time factor and model size highlights EnGen-TTS as a compelling choice, with efficient performance. This research marks a significant advancement in Bahasa TTS technology, with implications for diverse language applications. Link to Generated Samples: \url{https://bahasa-harmony-comp.vercel.app/}

语音合成印尼语数据集离散编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。