arXiv:2509.20086cs.CL2025-09被引 1

OLaPh提升多语言音素转换准确率,尤其擅长处理生词。

OLaPh: Optimal Language Phonemizer

  • 融合词典与神经网络,用统计子词分割提升泛化能力。
  • 在WikiPron上准确率超越现有方法,生词识别效果显著。
  • 适合语音合成、多语言研究者使用,开源可复现。

音素化是文本转语音合成中的关键环节。传统方法依赖确定性转换规则和词典,而神经方法在未登录词(OOV)上具有更好泛化能力。我们提出OLaPh(Optimal Language Phonemizer),一种结合大规模多语言词典、先进NLP技术与统计子词分割的混合框架。在WikiPron基准测试中,OLaPh在整体准确率上显著优于已有基线,并通过先进的回退机制在未登录词上保持鲁棒性。为进一步探索神经模型的泛化能力,我们利用该框架构建了一个高一致性训练语料,用于指令微调大型语言模型(LLM)。尽管确定性框架整体更准确,但该LLM展现出强大泛化能力,其性能达到甚至部分超过框架表现,表明其成功内化了合成数据中超越框架本身的语音规律。这些工具共同构成一个全面、开源的多语言字符到音素转换(G2P)研究资源。

原文摘要 · Abstract (English)

Phonemization is a critical component in text-to-speech synthesis. Traditional approaches rely on deterministic transformations and lexica, while neural methods offer potential for higher generalization on out-of-vocabulary (OOV) terms. We introduce OLaPh (Optimal Language Phonemizer), a hybrid framework that integrates extensive multilingual lexica with advanced NLP techniques and a statistical subword segmentation function. Evaluations on the WikiPron benchmark show OLaPh significantly outperforms established baselines in overall accuracy and maintains robustness on OOV data through advanced fallback mechanisms. To further explore neural generalization, we utilize the framework to synthesize a high-consistency training corpus for an instruction-tuned Large Language Model (LLM). While the deterministic framework remains more accurate overall, the LLM demonstrates strong generalization, matching or partly exceeding the framework's performance. This suggests that the LLM successfully internalized phonetic intuitions from the synthetic data that transcend the framework's capabilities. Together, these tools provide a comprehensive, open-source resource for multilingual grapheme-to-phoneme conversion (G2P) research.

语音合成音素化多语言LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。