arXiv:2606.20179cs.CL2026-06

用语音监督训练的希伯来语音素转换模型,更贴近真实口语发音。

ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion

  • 通过语音自动生成音素标签,避免人工标注拼音的繁琐
  • 在多个基准上超越现有最佳方法,尤其在口语数据上表现优异
  • 适合做希伯来语语音合成与语音识别的研究者使用

现代希伯来语的音素转换(G2P)对文本转语音(TTS)等应用至关重要,但因该语言为辅音文字,元音几乎不标出,导致严重歧义。传统方法先预测元音符号(nikud)再转音素,但标注成本高、无法反映实际口语中的重音等特征。直接序列到序列预测音素又受限于数据少,且难以利用辅音文字的字符级对齐特性。本文提出ReNikud,利用数千小时未标注希伯来语音频,通过基于音素的自动语音识别(ASR)伪标签生成反映自然口语的音素转录;设计伪声调架构,在每个字符位置预测音素,强制字符级对齐作为归纳偏置。在现有希伯来语G2P基准和新提出的针对口语的MILIM基准上,ReNikud均优于以往最优方法。代码与训练模型将公开,支持后续希伯来语语音技术研究。

原文摘要 · Abstract (English)

Grapheme-to-phoneme (G2P) conversion for Modern Hebrew is needed for applications like text-to-speech (TTS), but is challenging due to the language's abjad writing system, which leaves vowels largely unwritten, creating substantial ambiguity. Standard approaches first predict vowel diacritics (nikud) to produce International Phonetic Alphabet (IPA) transcriptions, but this is limited: vocalization data is scarce and laborious to produce, it does not specify features such as lexical stress, and it reflects formal grammatical rules rather than everyday spoken pronunciation. Direct sequence-to-sequence IPA prediction, meanwhile, struggles on limited data and fails to exploit the character-level alignment characteristic of abjads. Our method, ReNikud, overcomes these limitations with two key insights: (1) Weak audio supervision via a phoneme-based automatic speech recognition (ASR) pseudo-labeling pipeline on thousands of hours of unlabeled Hebrew audio, yielding phonemic transcriptions that reflect natural spoken norms without manual annotation. (2) A pseudo-vocalization architecture that predicts IPA phonemes at each character position, enforcing character-level alignment as an inductive bias. Results on existing Hebrew G2P benchmarks and the new targeted MILIM benchmark for spoken Hebrew show that ReNikud surpasses previous state-of-the-art methods. We will release our code and trained models to support further work on Hebrew TTS and speech technologies.

语音合成音素转换语音识别希伯来语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。