arXiv:2409.13262cs.CLcs.SD2024-09被引 12

用拼音增强大模型纠错,提升中文语音识别准确率

Large Language Model Should Understand Pinyin for Chinese ASR Error Correction

  • 引入拼音作为语音识别纠错的辅助信息
  • 在Aishell-1和Common Voice上优于纯文本方法
  • 通过多任务训练对齐拼音与文本特征空间

大型语言模型可通过生成式纠错增强自动语音识别系统。本文提出拼音增强型生成式纠错(PY-GEC),利用汉语拼音这一语音表示作为补充信息,提升中文语音识别错误纠正效果。该方法仅使用合成错误进行训练,并在推理时采用最优候选结果。此外,提出一种多任务训练策略,包含拼音与文本之间的转换任务,以对齐二者特征空间。在Aishell-1和Common Voice数据集上的实验表明,该方法始终优于仅使用文本输入的生成式纠错模型。更重要的是,从两个方面提供了有效性解释:1)拼音特征获得更高的注意力权重;2)拼音与文本隐藏状态的特征空间实现对齐。

原文摘要 · Abstract (English)

Large language models can enhance automatic speech recognition systems through generative error correction. In this paper, we propose Pinyin-enhanced GEC, which leverages Pinyi, the phonetic representation of Mandarin Chinese, as supplementary information to improve Chinese ASR error correction. Our approach only utilizes synthetic errors for training and employs the one-best hypothesis during inference. Additionally, we introduce a multitask training approach involving conversion tasks between Pinyin and text to align their feature spaces. Experiments on the Aishell-1 and the Common Voice datasets demonstrate that our approach consistently outperforms GEC with text-only input. More importantly, we provide intuitive explanations for the effectiveness of PY-GEC and multitask training from two aspects: 1) increased attention weight on Pinyin features; and 2) aligned feature space between Pinyin and text hidden states.

语音识别拼音增强生成纠错多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。