arXiv:2505.06599cs.CL2025-05

针对波斯语同形异音词难题,提出中间语言模型提升语音转换准确率。

Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation

  • 设计专用中间语言,结合大模型提示与序列到序列翻译架构。
  • 在双数据集上训练,显著降低音素错误率(PER),达新基准。
  • 适用于同形异音词多的语言,如中文、阿拉伯语,具可扩展性。

波斯语的音素转写(G2P)面临独特挑战,尤其源于其复杂的语音特征,包括同形异音词和不同语体中的“艾扎费”结构。本文提出一种专为波斯语设计的中间语言,通过多层面方法解决上述问题。方法结合大语言模型(LLM)提示技术与定制化的序列到序列机器转写架构。我们系统构建了一个包含多音字(即多音词)的完整词汇数据库,采用形式概念分析实现语义区分。模型在两个数据集上训练:由大模型生成的正式与非正式波斯语文本,以及B-Plus播客数据集中的非正式变体。实验结果表明,该模型在处理波斯语复杂音素转换方面优于现有最先进方法,显著降低音素错误率(PER),建立新基准。本研究推动低资源语言处理发展,为波斯语文本转语音系统提供可靠解决方案,并证明该方法可扩展至汉语、阿拉伯语等存在丰富同形异音现象的语言。

原文摘要 · Abstract (English)

Grapheme-to-phoneme (G2P) conversion for Persian presents unique challenges due to its complex phonological features, particularly homographs and Ezafe, which exist in formal and informal language contexts. This paper introduces an intermediate language specifically designed for Persian language processing that addresses these challenges through a multi-faceted approach. Our methodology combines two key components: Large Language Model (LLM) prompting techniques and a specialized sequence-to-sequence machine transliteration architecture. We developed and implemented a systematic approach for constructing a comprehensive lexical database for homographs with multiple pronunciations disambiguation often termed polyphones, utilizing formal concept analysis for semantic differentiation. We train our model using two distinct datasets: the LLM-generated dataset for formal and informal Persian and the B-Plus podcasts for informal language variants. The experimental results demonstrate superior performance compared to existing state-of-the-art approaches, particularly in handling the complexities of Persian phoneme conversion. Our model significantly improves Phoneme Error Rate (PER) metrics, establishing a new benchmark for Persian G2P conversion accuracy. This work contributes to the growing research in low-resource language processing and provides a robust solution for Persian text-to-speech systems and demonstrating its applicability beyond Persian. Specifically, the approach can extend to languages with rich homographic phenomena such as Chinese and Arabic

语音合成同形异音波斯语大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。