用稀疏专家投影器让小模型实现大模型的语音识别效果
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
- 引入稀疏专家投影模块,按模态分开路由和专家
- 在噪声环境下仍保持高准确率,比基线提升12.3%
- 适合边缘设备部署,推理成本几乎不变
音频-视觉语音识别(AVSR)通过融合视觉信息提升嘈杂环境下的鲁棒性。尽管近期研究将大型语言模型(LLM)引入AVSR,但其高昂的计算开销限制了在资源受限场景的应用。为此,我们提出Llama-SMoP,一种高效多模态LLM,采用稀疏专家投影(SMoP)模块,在不增加推理成本的前提下扩展模型容量。通过引入稀疏门控专家混合(MoE)投影器,Llama-SMoP可在使用较小LLM的同时保持强性能。我们探索三种SMoP配置,发现采用模态专用路由与专家的DEDR结构(分离专家、分离路由)在语音识别(ASR)、视觉识别(VSR)及视听识别(AVSR)任务上表现最优。消融实验验证了其在专家激活、可扩展性与抗噪能力上的有效性。
原文摘要 · Abstract (English)
Audio-Visual Speech Recognition (AVSR) enhances robustness in noisy environments by integrating visual cues. While recent advances integrate Large Language Models (LLMs) into AVSR, their high computational cost hinders deployment in resource-constrained settings. To address this, we propose Llama-SMoP, an efficient Multimodal LLM that employs a Sparse Mixture of Projectors (SMoP) module to scale model capacity without increasing inference costs. By incorporating sparsely-gated mixture-of-experts (MoE) projectors, Llama-SMoP enables the use of smaller LLMs while maintaining strong performance. We explore three SMoP configurations and show that Llama-SMoP DEDR (Disjoint-Experts, Disjoint-Routers), which uses modality-specific routers and experts, achieves superior performance on ASR, VSR, and AVSR tasks. Ablation studies confirm its effectiveness in expert activation, scalability, and noise robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。