arXiv:2604.12398eess.AS2026-04

用常见词的发音线索提升语音大模型对罕见词的识别准确率

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction

论文配图:Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction
图 1 · 摘自论文原文
  • 通过常见词的声学特征引导罕见词识别
  • 在跨领域数据上降低16.3%的罕见词识别错误
  • 无需音标知识或G2P工具,适合普通用户使用

语音感知大模型(SLLMs)近期实现了最先进的自动语音识别性能;然而,它们在转录训练数据中罕见或未出现过的偏置词时仍存在困难。现有上下文偏置机制通常通过文本提示或附加模块引入预定义的偏置词列表。为进一步提升性能,可将偏置词与其音素表示配对作为发音提示。通常音素序列通过覆盖目标语言和领域的音素到文本(G2P)系统生成。因此,当缺乏兼容的G2P系统时,基于音素的上下文偏置难以实施。此外,手动添加准确的音素序列需要高级语音学知识。本文提出一种基于与目标偏置词发音部分相似的常见词声学线索的上下文偏置方法。假设应用场景中终端用户无需掌握语音学知识,也不使用G2P工具进行推理。为增强鲁棒性,还引入多输出学习方式的偏置词位置预测。相比基线系统,该方法在罕见词识别错误上降低了16.3%,且在跨领域数据上表现良好。

原文摘要 · Abstract (English)

Speech-aware LLMs (SLLMs) have recently achieved state-of-the-art ASR performance; however, they still fail to accurately transcribe bias words that appear rarely or never in the training data. Contextual biasing mechanisms are commonly implemented by introducing a predefined bias word list into the model via a text prompt or additional module. For further improvement, predefined bias words can be paired with their phoneme representations as pronunciation cues. Typically, phoneme sequences are generated through a G2P system that covers the target languages and domains of the bias words. Therefore, when a compatible G2P system is unavailable, phoneme-assisted contextual biasing becomes difficult to perform. Moreover, manually adding accurate phoneme sequences requires advanced phonetic knowledge. In this paper, we explore contextual biasing in SLLM based on acoustic cues associated with a set of common words whose pronunciations are partially similar to those of the target bias words. We assume ASR applications in which end users do not require special knowledge of phonetics or utilize G2P tools for inference. For enhanced robustness, we also introduce bias word positional prediction implemented in a multi-output learning fashion. Our method reduces bias word recognition errors by 16.3% compared to baseline systems, including on out-of-domain data.

语音识别大模型罕见词声学线索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。