arXiv:2411.07563cs.AI2024-11中稿 · ISCSLP 2024被引 5

用大模型上下文检索提升语音转换中拼写到发音的准确性

Improving Grapheme-to-Phoneme Conversion through In-Context Knowledge Retrieval with Large Language Models

  • 利用大模型的上下文知识检索能力,解决拼写对应发音的歧义问题
  • 在LibriG2P数据集上,错误率降低2.0%(相对降幅28.9%)
  • 适合语音合成、自然语言处理领域研究者参考

拼写到发音(G2P)转换是文本转语音(TTS)系统中的关键步骤,负责将拼写映射为对应的音素表示。然而,该过程面临歧义问题:同一拼写在不同语境下可能对应多个音素,给G2P转换带来挑战。受大语言模型(LLM)在上下文感知任务中表现优异的启发,本文提出基于大模型上下文知识检索(ICKR)的上下文G2P转换系统,以增强消歧能力。在LibriG2P数据集上的实验表明,引入ICKR的最优系统相较基线,加权平均音素错误率(PER)绝对降低2.0%(相对降低28.9%)。使用GPT-4作为ICKR组件时,错误率进一步降低3.5%(相对降低3.8%)。

原文摘要 · Abstract (English)

Grapheme-to-phoneme (G2P) conversion is a crucial step in Text-to-Speech (TTS) systems, responsible for mapping grapheme to corresponding phonetic representations. However, it faces ambiguities problems where the same grapheme can represent multiple phonemes depending on contexts, posing a challenge for G2P conversion. Inspired by the remarkable success of Large Language Models (LLMs) in handling context-aware scenarios, contextual G2P conversion systems with LLMs' in-context knowledge retrieval (ICKR) capabilities are proposed to promote disambiguation capability. The efficacy of incorporating ICKR into G2P conversion systems is demonstrated thoroughly on the Librig2p dataset. In particular, the best contextual G2P conversion system using ICKR outperforms the baseline with weighted average phoneme error rate (PER) reductions of 2.0% absolute (28.9% relative). Using GPT-4 in the ICKR system can increase of 3.5% absolute (3.8% relative) on the Librig2p dataset.

语音合成大模型G2P上下文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。