arXiv:2511.08389eess.AScs.AI2025-11中稿 · IEEE ASRU 2025

统一模型与层融合,提升语音大模型下游表现。

Unifying Model and Layer Fusion for Speech Foundation Models

  • 设计接口模块,实现多模型多层表示的联合融合。
  • 在语音识别与副语言分析任务中优于现有融合方法。
  • 适合希望最大化利用多个语音大模型性能的研究者。

语音基础模型近年来受到广泛关注。先前研究已表明,融合同一模型的多层表征或融合多个模型的输出,可提升下游任务性能。本文提出一种统一接口模块,实现跨多个上游语音模型的融合,并整合各模型内部的多层信息。我们在多种自监督与监督模型上,针对语音识别和副语言分析等任务进行了广泛实验,结果表明该方法优于现有融合策略。进一步分析了模型规模与数量对可扩展性的影响,强调了选择合适上游模型的重要性。实验显示,在合理选择上游模型的前提下,所提接口可带来额外性能提升,是利用语音基础模型的有前景方案。

原文摘要 · Abstract (English)

Speech Foundation Models have gained significant attention recently. Prior works have shown that the fusion of representations from multiple layers of the same model or the fusion of multiple models can improve performance on downstream tasks. We unify these two fusion strategies by proposing an interface module that enables fusion across multiple upstream speech models while integrating information across their layers. We conduct extensive experiments on different self-supervised and supervised models across various speech tasks, including ASR and paralinguistic analysis, and demonstrate that our method outperforms prior fusion approaches. We further analyze its scalability concerning model size and count, highlighting the importance of selecting appropriate upstream models. Our results show that the proposed interface provides an additional performance boost when given a suitable upstream model selection, making it a promising approach for utilizing Speech Foundation Models.

语音模型模型融合基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。