用Mamba架构提升多语言语音识别的效率与准确性
MLMA: Towards Multilingual ASR With Mamba-based Architectures
- 采用Mamba状态空间模型处理长序列,替代传统Transformer
- 在多语言基准上表现媲美Transformer,兼顾高低资源语言
- 适合需要高效多语种语音识别的场景,如跨语言服务
多语言自动语音识别(ASR)仍具挑战性,尤其在高资源与低资源语言间平衡性能时。近期序列建模进展表明,超越Transformer的架构可能具备更好的可扩展性与效率。本文提出MLMA(基于Mamba的多语言语音建模),利用Mamba——一种专为长上下文序列处理优化的状态空间模型——实现多语言ASR。MLMA通过隐式语言感知条件与共享表征,在多种语言间实现稳健识别。在标准多语言基准上的实验表明,其性能可与基于Transformer的模型相媲美。结果凸显Mamba作为可扩展、高效且精准的多语言语音识别骨干网络的巨大潜力。
原文摘要 · Abstract (English)
Multilingual automatic speech recognition (ASR) remains a challenging task, especially when balancing performance across high- and low-resource languages. Recent advances in sequence modeling suggest that architectures beyond Transformers may offer better scalability and efficiency. In this work, we introduce MLMA (Multilingual Language Modeling with Mamba for ASR), a new approach that leverages the Mamba architecture -- an efficient state-space model optimized for long-context sequence processing -- for multilingual ASR. Using Mamba, MLMA implicitly incorporates language-aware conditioning and shared representations to support robust recognition across diverse languages. Experiments on standard multilingual benchmarks show that MLMA achieves competitive performance compared to Transformer-based architectures. These results highlight Mamba's potential as a strong backbone for scalable, efficient, and accurate multilingual speech recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。