对比两款日语角色语音合成模型,发现新模型更自然、更准确。
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
- 用音调强调控制和声学判别器提升语音表现力
- 新模型自然度达4.37分(接近真人4.38分),错误率更低
- 适合语言学习与角色对话生成,但计算开销较大
由于日语音调重音敏感性和风格多变性,合成具有表现力的日语角色语音面临独特挑战。本文针对领域内、角色驱动的日语语音,实证评估了两种开源语音合成模型——VITS与Style-BERT-VITS2 JP Extra(SBV2JE)。基于三个角色专属数据集,从自然度(平均意见评分与对比评分)、可懂度(词错误率)及说话人一致性三方面进行评估。结果表明,SBV2JE在自然度上接近真人水平(MOS 4.37 vs. 4.38),词错误率更低,对比评分略有偏好。得益于音调强调控制与基于WavLM的判别器,该模型在语言学习与角色对话语音生成中表现优异,尽管计算成本更高。
原文摘要 · Abstract (English)
Synthesizing expressive Japanese character speech poses unique challenges due to pitch-accent sensitivity and stylistic variability. This paper empirically evaluates two open-source text-to-speech models--VITS and Style-BERT-VITS2 JP Extra (SBV2JE)--on in-domain, character-driven Japanese speech. Using three character-specific datasets, we evaluate models across naturalness (mean opinion and comparative mean opinion score), intelligibility (word error rate), and speaker consistency. SBV2JE matches human ground truth in naturalness (MOS 4.37 vs. 4.38), achieves lower WER, and shows slight preference in CMOS. Enhanced by pitch-accent controls and a WavLM-based discriminator, SBV2JE proves effective for applications like language learning and character dialogue generation, despite higher computational demands.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。