Mamba在语音重建任务中表现优异,但识别任务需额外模块支持。
Rethinking Mamba in Speech Processing by Self-Supervised Models
- 基于信息论分析,Mamba擅长语音重建而非直接分类。
- 构建Mamba-HuBERT模型,验证了重建步骤对识别的关键作用。
- 适合研究语音建模中序列架构与任务匹配关系的学者。
基于Mamba的模型在计算机视觉、自然语言处理及语音处理中均表现出色。然而,在语音处理领域,其性能因任务而异:在语音增强和频谱重建等重建类任务中表现良好,但在语音识别等分类任务中,需引入额外模块才能超越注意力模型。本文提出假设:Mamba在语音重建任务中具有优势,而语音识别等分类任务需先完成重建步骤。通过信息论视角分析已有Mamba语音模型,并结合HuBERT特性,训练了基于Mamba的HuBERT模型。实验结果表明,模型的互信息模式与性能指标均支持该假设。
原文摘要 · Abstract (English)
The Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model's performance varies across different tasks. For instance, in tasks such as speech enhancement and spectrum reconstruction, the Mamba model performs well when used independently. However, for tasks like speech recognition, additional modules are required to surpass the performance of attention-based models. We propose the hypothesis that the Mamba-based model excels in "reconstruction" tasks within speech processing. However, for "classification tasks" such as Speech Recognition, additional modules are necessary to accomplish the "reconstruction" step. To validate our hypothesis, we analyze the previous Mamba-based Speech Models from an information theory perspective. Furthermore, we leveraged the properties of HuBERT in our study. We trained a Mamba-based HuBERT model, and the mutual information patterns, along with the model's performance metrics, confirmed our assumptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。