arXiv:2509.12647cs.CLeess.AS2025-09被引 3

让语音识别更懂发音,精准区分同音词。

PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition

  • 用字形与音素交叉建模,引导模型关注发音线索。
  • 在英文和中文数据集上,错误率降低超30%以上。
  • 特别擅长识别生僻词,适合语音识别研究者使用。

本文提出一种发音感知的上下文大语言模型框架(PAC),解决基于大语言模型的自动语音识别系统中发音建模与同音词区分两大难题,二者对原始词或长尾词识别至关重要。该方法采用两阶段学习范式:首先引入发音引导的上下文学习策略,通过交错的音素-字素上下文建模并加入仅含字素的干扰项,促使模型利用音素信息实现准确识别;其次提出带扰动标签采样的发音区分性强化学习方法,进一步提升模型对上下文同音词的辨别能力。在公开的英文LibriSpeech和中文AISHELL-1数据集上的实验表明,PAC相较预训练大语言模型基线系统,分别将相对词错误率(WER)降低30.2%和53.8%;对长尾词的偏差词错误率分别减少31.8%和60.5%。

原文摘要 · Abstract (English)

This paper presents a Pronunciation-Aware Contextualized (PAC) framework to address two key challenges in Large Language Model (LLM)-based Automatic Speech Recognition (ASR) systems: effective pronunciation modeling and robust homophone discrimination. Both are essential for raw or long-tail word recognition. The proposed approach adopts a two-stage learning paradigm. First, we introduce a pronunciation-guided context learning method. It employs an interleaved grapheme-phoneme context modeling strategy that incorporates grapheme-only distractors, encouraging the model to leverage phonemic cues for accurate recognition. Then, we propose a pronunciation-discriminative reinforcement learning method with perturbed label sampling to further enhance the modelś ability to distinguish contextualized homophones. Experimental results on the public English Librispeech and Mandarin AISHELL-1 datasets indicate that PAC: (1) reduces relative Word Error Rate (WER) by 30.2% and 53.8% compared to pre-trained LLM-based ASR models, and (2) achieves 31.8% and 60.5% relative reductions in biased WER for long-tail words compared to strong baselines, respectively.

语音识别大模型发音建模同音词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。