arXiv:2508.08585eess.AS2025-08被引 3

通过联合解码控制语音识别中的上下文信息注入程度

Joint decoding method for controllable contextual speech recognition based on Speech LLM

  • 提出联合解码方法,显式调控上下文信息注入强度
  • 在未预训练长上下文数据的模型上实现长上下文理解能力
  • 可同时用于敏感词抑制和可控语音识别

上下文语音识别指根据上下文信息识别特定内容偏好。近期,利用语音大模型(Speech LLM)的上下文理解能力,通过提示注入上下文信息已成为研究热点。然而,直接通过提示注入的方法依赖模型内部注意力机制,无法显式控制信息注入程度。为此,本文提出一种联合解码方法,实现对注入上下文信息的显式控制,并取得更优的识别性能。此外,该方法还可用于敏感词抑制识别。实验表明,即使模型未在长上下文数据上预训练,也能通过本方法获得长上下文理解能力。

原文摘要 · Abstract (English)

Contextual speech recognition refers to the ability to identify preferences for specific content based on contextual information. Recently, leveraging the contextual understanding capabilities of Speech LLM to achieve contextual biasing by injecting contextual information through prompts have emerged as a research hotspot.However, the direct information injection method via prompts relies on the internal attention mechanism of the model, making it impossible to explicitly control the extent of information injection. To address this limitation, we propose a joint decoding method to control the contextual information. This approach enables explicit control over the injected contextual information and achieving superior recognition performance. Additionally, Our method can also be used for sensitive word suppression recognition.Furthermore, experimental results show that even Speech LLM not pre-trained on long contextual data can acquire long contextual capabilities through our method.

语音识别上下文控制大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。