解决日语汉字多音字问题,提升日语语音合成准确率
Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis

- 通过扩大数据规模并针对2136个常用汉字设计增强数据
- 在汉字级发音准确率上达到当前最佳,零样本合成效果最优
- 专为日语设计的评测基准与读音评估指标,适合语音合成研究者
尽管基于大语言模型(LLM)的文本转语音(TTS)系统已实现高质量语音合成,但现有系统多聚焦于英语和中文。日语仍鲜被深入研究,其独特语言挑战如广泛存在的上下文相关汉字多音现象尚未有效解决。本文提出Sarashina2.2-TTS(https://github.com/sbintuitions/sarashina2.2-tts),一种以日语为中心的LLM-TTS系统,采用双轨策略应对挑战:数据策略与评估方法。首先,将训练数据扩展至约361,000小时,包含均衡的日语与英语数据;同时设计针对性的数据增强流程,覆盖日本文化厅规定的全部2,136个常用汉字(Joyo kanji),高效解决汉字多音歧义问题。其次,提出Joyo Kanji Yomi基准(https://github.com/sbintuitions/JoyoKanji-Yomi-Benchmark),涵盖全部2,136个常用汉字及其4,378种读音;并引入Kana-CER指标,通过假名空间比较合成语音与标准读音,消除拼写差异,直接衡量发音正确性。实验表明,目标数据增强显著提升发音准确率。总体而言,Sarashina2.2-TTS在汉字级发音准确率上达到领先水平,句子级发音表现媲美顶级基线,且在零样本日语语音合成中实现最高说话人相似度。跨语言评估进一步证实,该系统在不同提示语言下仍能保持稳定的日语发音表现,验证了平衡训练对跨语言鲁棒性的提升作用。
原文摘要 · Abstract (English)
While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chinese. Japanese, however, remains under-explored, and its unique linguistic challenges, such as widespread context-dependent kanji polyphony, have yet to be adequately tackled. Here we introduce Sarashina2.2-TTS (https://github.com/sbintuitions/sarashina2.2-tts), a Japanese-centric LLM-TTS system that tackles these challenges through a dual approach: data strategy and evaluation methodology. First, we scale training to approximately 361k hours of speech, incorporating a balanced mix of Japanese and English data. Furthermore, we design a targeted data augmentation pipeline covering all 2,136 Joyo (regular-use) kanji designated by Japan's Agency for Cultural Affairs to efficiently address kanji polyphony disambiguation. Second, we introduce the Joyo Kanji Yomi Benchmark (https://github.com/sbintuitions/JoyoKanji-Yomi-Benchmark), covering all 2,136 Joyo kanji and their 4,378 readings. Alongside this benchmark, we propose Kana-CER, a metric that compares synthesized speech against reference readings in the kana space, eliminating orthographic variations to directly measure pronunciation correctness. Experiments demonstrate that our targeted data augmentation significantly improves reading accuracy. Overall, Sarashina2.2-TTS achieves state-of-the-art kanji-level reading accuracy and matches top baselines on general sentence-level pronunciation, while delivering the highest speaker similarity in zero-shot Japanese speech synthesis. Furthermore, cross-lingual evaluation reveals that Sarashina2.2-TTS is the only system that maintains stable Japanese pronunciation regardless of the prompt language, confirming that our balanced training approach improves cross-lingual robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。