解决希伯来语语音合成中的发音不完整问题,提升语音自然度。
Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech
- 基于字形转音素系统,生成完整国际音标标注。
- 新数据集与基准测试验证模型优于已有方法。
- 小模型结合该系统可媲美大型商业系统,适合资源有限者使用。
现代希伯来语的语音合成面临拼写复杂带来的挑战,现有方法常忽略如重音等未明确标注的发音特征。本文提出一个框架以实现更准确的希伯来语语音合成,包含四项贡献:(1) Phonikud,一个开源的希伯来语字形转音素(G2P)系统,能输出完整的国际音标(IPA)转写,通过增强基础带调标记器构建;(2) ILSpeech语料库,包含配对的希伯来语文本、音频及专家标注的IPA;(3) 首个针对希伯来语G2P转换任务的基准测试;(4) 能捕捉以往被忽视的发音细节的音频到IPA模型,用于自动语音合成评估。实验表明,Phonikud在预测希伯来语音素方面优于先前方法,且使用Phonikud输入的小型本地化语音合成模型可达到大型专有系统的水平。代码、数据和模型已公开于 https://phonikud.github.io。
原文摘要 · Abstract (English)
Text-to-speech (TTS) for Modern Hebrew is challenged by the language's orthographic complexity, with existing solutions ignoring underspecified phonetic features such as stress. We present a framework for more phonetically accurate Hebrew TTS with four contributions: (1) Phonikud, an open-source Hebrew grapheme-to-phoneme (G2P) system that outputs fully-specified International Phonetic Alphabet (IPA) transcriptions, designed by augmenting a base diacritizer. (2) The ILSpeech corpus of paired Hebrew audio, text, and expert IPA annotations. (3) A benchmark for the previously unmeasured task of Hebrew G2P conversion. (4) Hebrew audio-to-IPA models capturing previously disregarded phonetic details for automatic TTS evaluation. Our results show that Phonikud more accurately predicts Hebrew phonemes than prior methods, and that small, local TTS models with phonetic input from Phonikud approach large proprietary systems. We release our code, data, and models at https://phonikud.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。