arXiv:2412.10989eess.AScs.SD2024-12被引 4

用Mamba提升语音验证的长序列建模能力与效率

MASV: Speaker Verification with Global and Local Context Mamba

  • 在ECAPA-TDNN中引入局部双向Mamba和三重Mamba模块
  • 在SPEAKERVERIFICATION19等数据集上准确率显著提升
  • 兼顾长序列建模与计算效率,适合实际部署

卷积神经网络和变换器等深度学习模型在语音验证中表现优异,但基于CNN的方法难以有效建模长序列音频,导致验证性能受限;而基于变换器的方法则因高计算开销限制了实用性。本文提出MASV模型,将Mamba模块融入ECAPA-TDNN框架,通过引入局部上下文双向Mamba和三重Mamba模块,有效捕捉音频序列中的全局与局部上下文信息。实验表明,MASV模型在准确率和效率方面均显著优于现有方法,在SPEAKERVERIFICATION19、VoxCeleb1等数据集上实现更高性能。

原文摘要 · Abstract (English)

Deep learning models like Convolutional Neural Networks and transformers have shown impressive capabilities in speech verification, gaining considerable attention in the research community. However, CNN-based approaches struggle with modeling long-sequence audio effectively, resulting in suboptimal verification performance. On the other hand, transformer-based methods are often hindered by high computational demands, limiting their practicality. This paper presents the MASV model, a novel architecture that integrates the Mamba module into the ECAPA-TDNN framework. By introducing the Local Context Bidirectional Mamba and Tri-Mamba block, the model effectively captures both global and local context within audio sequences. Experimental results demonstrate that the MASV model substantially enhances verification performance, surpassing existing models in both accuracy and efficiency.

语音验证Mamba序列建模高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。