arXiv:2505.18860eess.AS2025-05中稿 · Interspeech 2025被引 3

根据语音上下文动态剪枝,显著降低计算量并提升识别效果

Context-Driven Dynamic Pruning for Large Speech Foundation Models

  • 依据说话人、语言和声学事件等上下文动态调整模型计算
  • 推理时减少56.7 GFLOPs计算量,BLEU分数提升25.7%
  • 适合资源受限场景下的语音大模型高效部署

语音基础模型在多语言和多声学条件下表现出强泛化能力,但推理时需要大量计算资源。现有剪枝方法通过利用外部上下文动态优化模型结构。本文提出上下文驱动的动态剪枝技术,根据输入帧间的上下文关系及额外信息在推理时优化模型计算。采用Open Whisper-style Speech Model (OWSM),引入说话人嵌入、声学事件嵌入和语言信息作为上下文。实验表明,结合说话人嵌入后,模型推理计算量减少56.7 GFLOPs,同时相比全微调的OWSM模型,BLEU分数相对提升25.7%。

原文摘要 · Abstract (English)

Speech foundation models achieve strong generalization across languages and acoustic conditions, but require significant computational resources for inference. In the context of speech foundation models, pruning techniques have been studied that dynamically optimize model structures based on the target audio leveraging external context. In this work, we extend this line of research and propose context-driven dynamic pruning, a technique that optimizes the model computation depending on the context between different input frames and additional context during inference. We employ the Open Whisper-style Speech Model (OWSM) and incorporate speaker embeddings, acoustic event embeddings, and language information as additional context. By incorporating the speaker embedding, our method achieves a reduction of 56.7 GFLOPs while improving BLEU scores by a relative 25.7% compared to the fully fine-tuned OWSM model.

语音模型动态剪枝高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。