arXiv:2603.27205cs.SD2026-03

让大模型持续访问说话人分离的声学记忆,提升三人群体语音识别性能。

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition

  • 用可检索的声学记忆替代初始声学前缀,实现全程说话人信息追踪。
  • 在三人群体语音识别中相对基线提升2.7% WER,噪声环境下仍稳定增益。
  • 适合需要高精度多人语音识别的场景,如会议记录、听障辅助系统。

大型语言模型(LLM)在双人语音识别中表现优异,但在三人混合语音场景下性能显著下降。核心问题在于传统系统仅通过初始投影前缀提供声学证据,迫使解码器在自回归生成过程中持续保持精细的说话人信息。本文重新审视基于CTC的静态前缀条件化,包括离散标记、混合标记-声学与连续声学提示。尽管连续声学提示更可靠,但在Libri3Mix上的改进仍有限,说明丰富前缀内容无法突破条件瓶颈。为此,提出序列化声学记忆中的持续定位机制,将说话人分离且按起始时间排序的声学表示作为外部记忆,在解码时通过门控残差交叉注意力访问。进一步引入低秩联合优化方法(LoRA),对声学检索路径与选定的LLM自注意力投影进行协同微调。在干净与噪声条件下的Libri2Mix和Libri3Mix实验显示,相比传统LLM-SOT与简单堆叠交叉注意力,本方法始终取得一致提升,尤其在三人混合场景中优势显著。结果表明,持续获取结构化、说话人感知的声学证据对基于大模型的多说话人语音识别至关重要。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are effective decoders for Serialized Output Training (SOT) in two-talker automatic speech recognition (ASR), but their performance degrades substantially in three-talker mixtures. A key limitation is that conventional systems provide acoustic evidence only through an initial projected prefix, requiring the decoder to preserve fine-grained talker information throughout autoregressive generation. We first revisit CTC-derived static prefix conditioning using discrete token, hybrid token-acoustic, and continuous acoustic prompts. Although continuous acoustic cues are more reliable than discrete CTC hypotheses, the improvements on Libri3Mix remain limited, showing that richer prefix content alone does not resolve the conditioning bottleneck. We therefore propose persistent grounding in serialized acoustic memory, which enables the decoder to retrieve talker-structured acoustic evidence throughout the utterance. Specifically, talker-disentangled and onset-ordered acoustic representations are retained as external memory and accessed during decoding through gated residual cross-attention. We further introduce joint low-rank refinement of the acoustic retrieval pathway and selected LLM self-attention projections using LoRA. Experiments on Libri2Mix and Libri3Mix under clean and noisy conditions show consistent improvements over conventional LLM-SOT and naive stacked cross-attention, with particularly large gains in three-talker mixtures. These results demonstrate the importance of persistent access to structured, talker-aware acoustic evidence for LLM-based multi-talker ASR.

语音识别多说话人大模型声学记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。