arXiv:2508.01166cs.SD2025-08AAAI被引 4

用多模态检索选择关键历史信息,让语音识别更准更快

Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASR

  • 通过声学与文本相似性筛选最相关历史对话
  • 仅用1.5小时数据训练效果超越用179小时数据的顶尖系统
  • 适合资源有限但需高精度对话语音识别的场景

自动语音识别(ASR)旨在将人类语音内容转化为对应文本。在对话场景中,有效利用上下文可提升识别准确率。大语言模型(LLM)具备出色的长上下文理解与推理能力,使基于LLM的语音识别(LLM-ASR)能够利用历史对话上下文识别具有高度语境相关性的口语。然而,现有方法通常固定使用前几条对话或全部历史记录作为上下文,导致大量无关冗余信息引入,造成显著识别混淆和计算开销。本文提出一种名为MARS的多模态检索-选择方法,通过检索并选择与当前话语最相关的声学与文本历史上下文,增强对话式LLM-ASR。具体而言,多模态检索获取一组候选历史上下文,其在声学或文本上与当前话语高度相似;多模态选择计算每个候选的历史上下文的声学与文本相似度,并采用所提出的近似理想排序方法综合考虑二者,选出最优上下文。在Interspeech 2025多语言对话语音语言模型挑战赛数据集上的评估表明,仅使用1.5K小时数据训练的LLM-ASR,在集成MARS后,性能超过使用179K小时数据训练的顶尖系统。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) aims to convert human speech content into corresponding text. In conversational scenarios, effectively utilizing context can enhance its accuracy. Large Language Models' (LLMs) exceptional long-context understanding and reasoning abilities enable LLM-based ASR (LLM-ASR) to leverage historical context for recognizing conversational speech, which has a high degree of contextual relevance. However, existing conversational LLM-ASR methods use a fixed number of preceding utterances or the entire conversation history as context, resulting in significant ASR confusion and computational costs due to massive irrelevant and redundant information. This paper proposes a multi-modal retrieval-and-selection method named MARS that augments conversational LLM-ASR by enabling it to retrieve and select the most relevant acoustic and textual historical context for the current utterance. Specifically, multi-modal retrieval obtains a set of candidate historical contexts, each exhibiting high acoustic or textual similarity to the current utterance. Multi-modal selection calculates the acoustic and textual similarities for each retrieved candidate historical context and, by employing our proposed near-ideal ranking method to consider both similarities, selects the best historical context. Evaluations on the Interspeech 2025 Multilingual Conversational Speech Language Model Challenge dataset show that the LLM-ASR, when trained on only 1.5K hours of data and equipped with the MARS, outperforms the state-of-the-art top-ranking system trained on 179K hours of data.

语音识别多模态对话系统LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。