arXiv:2410.00070eess.AScs.CL2024-10中稿 · ICASSP 2025被引 12

用Mamba提升中文流式语音识别速度与准确率

Mamba for Streaming ASR Combined with Unimodal Aggregation

  • 采用Mamba编码器结合可控前瞻机制,实现高效流式识别
  • 在两个中文数据集上达到高精度与低延迟的平衡表现
  • 提出动态触发输出与特征聚合方法,适合实时语音应用

本文研究流式自动语音识别(ASR)。近期提出的状态空间模型Mamba在多项任务中表现媲美甚至超越Transformer,且具有线性复杂度优势。我们探索了Mamba编码器在流式ASR中的效率,并提出一种可控未来信息利用的前瞻机制。此外,引入一种流式单模态聚合(UMA)方法,能自动检测词元活动并即时触发输出,同时聚合特征帧以增强词元表示学习。基于UMA,进一步提出早期终止(ET)方法以降低识别延迟。在两个普通话中文数据集上的实验表明,所提模型在识别准确率和延迟方面均达到竞争力表现。

原文摘要 · Abstract (English)

This paper works on streaming automatic speech recognition (ASR). Mamba, a recently proposed state space model, has demonstrated the ability to match or surpass Transformers in various tasks while benefiting from a linear complexity advantage. We explore the efficiency of Mamba encoder for streaming ASR and propose an associated lookahead mechanism for leveraging controllable future information. Additionally, a streaming-style unimodal aggregation (UMA) method is implemented, which automatically detects token activity and streamingly triggers token output, and meanwhile aggregates feature frames for better learning token representation. Based on UMA, an early termination (ET) method is proposed to further reduce recognition latency. Experiments conducted on two Mandarin Chinese datasets demonstrate that the proposed model achieves competitive ASR performance in terms of both recognition accuracy and latency.

流式识别Mamba语音识别低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。