用大模型提升拼音转换能力,无需训练就能超越传统工具。
LLM-Powered Grapheme-to-Phoneme Conversion: Benchmark and Case Study
- 用提示工程和后处理增强大模型输出,不需额外训练
- 在波斯语句子级语音挑战上表现优于传统方法
- 适合语音合成、低资源语言处理的研究者
音素转写(G2P)在语音处理中至关重要,尤其在语音合成等应用中。传统G2P系统需具备语言学理解力与上下文感知能力,以应对多音字和上下文依赖的音素。大型语言模型(LLMs)在多种语言任务中展现出巨大潜力,表明其可能蕴含可用于G2P的语音知识。本文评估了LLMs在G2P中的性能,并提出无需额外训练或标注数据的提示工程与后处理方法来提升输出质量。我们还构建了一个针对波斯语句级语音挑战的基准数据集。实验结果表明,在采用上述方法后,LLMs在波斯语这一低资源语言上的表现甚至超过了传统G2P工具,凸显了基于大模型的G2P系统的发展潜力。
原文摘要 · Abstract (English)
Grapheme-to-phoneme (G2P) conversion is critical in speech processing, particularly for applications like speech synthesis. G2P systems must possess linguistic understanding and contextual awareness of languages with polyphone words and context-dependent phonemes. Large language models (LLMs) have recently demonstrated significant potential in various language tasks, suggesting that their phonetic knowledge could be leveraged for G2P. In this paper, we evaluate the performance of LLMs in G2P conversion and introduce prompting and post-processing methods that enhance LLM outputs without additional training or labeled data. We also present a benchmarking dataset designed to assess G2P performance on sentence-level phonetic challenges of the Persian language. Our results show that by applying the proposed methods, LLMs can outperform traditional G2P tools, even in an underrepresented language like Persian, highlighting the potential of developing LLM-aided G2P systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。