arXiv:2606.00507cs.CL2026-06

用隐式推理提升语音识别的上下文理解能力

LaSR: Context-Aware Speech Recognition via Latent Reasoning

论文配图:LaSR: Context-Aware Speech Recognition via Latent Reasoning
图 1 · 摘自论文原文
  • 通过声学特征区域对齐思维链监督,实现隐式推理
  • 在学术术语识别上显著优于传统微调方法
  • 适合需要精准理解专业语境的语音助手场景

近年来,语音大语言模型在口语理解与推理方面取得显著进展,但其上下文感知能力仍受限,难以有效反映说话人意图和主题背景。本文提出LaSR(隐式语音推理)训练范式,通过隐式推理过程增强上下文感知。不同于生成显式中间标记,LaSR将思维链(CoT)监督对齐于目标词的声学特征区域,并引入隐式推理时段以实现上下文信息锚定与转录过渡。此外,为有效评估特定词汇的上下文识别能力,我们构建了聚焦学术术语的大规模语料Spoken Darwin-Science。Fun-Audio-Chat上的初步实验表明,LaSR在不增加额外延迟的情况下显著提升术语识别效果,且持续优于标准监督微调基线。研究结果表明,隐式推理在构建高效、上下文感知的语音助手方面具有巨大潜力。

原文摘要 · Abstract (English)

Recent advances in Speech Large Language Models (Speech LLMs) have significantly enhanced spoken language understanding and reasoning. However, their contextual awareness is limited, struggling to perform speech recognition that effectively reflects the speaker's intent and topical context. In this paper, we propose LaSR (Latent Speech Reasoning), a novel training paradigm featuring a context-aware reasoning trajectory that leverages the latent reasoning process. Instead of generating explicit intermediate tokens, LaSR aligns chain-of-thought (CoT) supervision around the acoustic feature region of the targeted word, and introduces latent reasoning periods for context information grounding and transcriptional transition. Furthermore, to effectively benchmark contextual recognition on specialized vocabulary, we propose Spoken Darwin-Science, a large-scale corpus focusing on academic terminologies. Preliminary experiments on Fun-Audio-Chat demonstrate that LaSR significantly improves terminology recognition without introducing additional latency and consistently outperforms standard supervised fine-tuning baselines. Our findings highlight the potential of latent reasoning in building efficient, context-aware speech assistants.

语音识别隐式推理上下文感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。