arXiv:2506.01322cs.CLcs.SD2025-06ACL被引 6

构建941小时越南语语音数据集,提升零样本语音合成效果

Zero-Shot Text-to-Speech for Vietnamese

  • 构建941小时高质量越南语语音数据集PhoAudiobook
  • VALL-E和VoiceCraft在短句合成上表现更优
  • 公开数据集推动越南语语音合成研究

本文提出PhoAudiobook,一个包含941小时高质量音频的越南语文本到语音数据集。基于该数据集,我们在三种主流零样本语音合成模型(VALL-E、VoiceCraft、XTTS-V2)上开展实验,结果表明PhoAudiobook能持续提升模型性能。此外,VALL-E与VoiceCraft在短句生成中表现更佳,展现出更强的跨语言适应能力。研究团队已公开发布PhoAudiobook,以促进越南语语音合成领域的进一步研究与发展。

原文摘要 · Abstract (English)

This paper introduces PhoAudiobook, a newly curated dataset comprising 941 hours of high-quality audio for Vietnamese text-to-speech. Using PhoAudiobook, we conduct experiments on three leading zero-shot TTS models: VALL-E, VoiceCraft, and XTTS-V2. Our findings demonstrate that PhoAudiobook consistently enhances model performance across various metrics. Moreover, VALL-E and VoiceCraft exhibit superior performance in synthesizing short sentences, highlighting their robustness in handling diverse linguistic contexts. We publicly release PhoAudiobook to facilitate further research and development in Vietnamese text-to-speech.

语音合成零样本越南语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。