arXiv:2606.10838eess.AS2026-06中稿 · Interspeech 2026

用视频描述引导语音模型推理,提升罕见词识别准确率

Towards Deep Contextual Reasoning from Broad Descriptions for ASR with Speech-LLM via Metadata-Driven Reasoning Chains

  • 以视频元数据为弱先验,构建链式推理机制
  • 在YouTube测试集上显著降低罕见词错误率
  • 适合需要上下文理解的语音识别场景

语音识别在罕见领域术语和上下文相关命名实体上表现不佳。现有上下文方法多依赖关键词或短语列表,难以扩展且未能挖掘深层知识。本文提出一种训练方法,让语音大模型利用视频等宽泛描述作为弱语义先验,基于音频进行上下文推理。通过将错误识别结果与视频元数据、大模型生成的推理解释配对,构建了400小时的推理增强语音数据。微调语音大模型实现思维链推理:先生成初始转录,再结合上下文推理,最终输出修正转录。在来自YouTube的独立测试集上,该方法有效减少错误,尤其在罕见词和命名实体上表现更优,为语音识别中的深度上下文推理奠定基础。

原文摘要 · Abstract (English)

Speech recognition often fails on rare, domain-specific terms and context-related named entities. Existing contextualization techniques typically bias decoding with keywords or phrase lists, which does not scale well or exploit deeper knowledge. We propose a training method that teaches a speech-LLM to use broad descriptions (e.g. from videos) as weak semantic priors to perform contextual reasoning grounded in the audio. We build 400 hours of reasoning-augmented speech data by pairing erroneous hypotheses with video metadata and LLM-generated reasoning explanations that justify context-driven corrections. We finetune the speech-LLM to perform chain-of-thought reasoning: generate an initial transcript, then reason over the context, and finally return a corrected transcript. On held-out YouTube-derived test sets, our approach reduces errors, with specific improvements on rare words and named entities, and lays groundwork for deeper contextual reasoning in speech recognition.

语音识别链式推理上下文理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。