arXiv:2506.21576cs.CLcs.AI2025-06中稿 · Interspeech 2025被引 4

用软提示微调让Whisper更好识别混语语音,省参数还不丢老技能。

Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning

  • 用可学习的软提示代替全模型微调,只改少量参数提升混语识别能力。
  • 在SEAME和ASRU2019数据集上,新方法使混语语音识别错误率进一步下降。
  • 适合资源少的语言或混语场景,对已有语言性能无影响,效率高。

大型多语言语音识别模型如Whisper在高资源环境下表现优异,但在低资源场景(如稀有语言和混语)中因计算成本高和灾难性遗忘而受限。本文探索了参数高效的软提示微调(SPT)方法,以增强混语语音识别并保留原有知识。评估了两种策略:(1) 全模型微调(FFT),同时优化软提示和整个Whisper模型,相比传统方法提升了跨语言能力;(2) 遵循SPT原始设计,冻结模型参数仅训练软提示。此外,提出SPT4ASR,融合多种SPT变体。在SEAME和ASRU2019数据集上的实验表明,深层提示微调是最有效的SPT方法,SPT4ASR进一步降低了混语识别错误率,参数效率与LoRA相当,且未损害原有语言性能。

原文摘要 · Abstract (English)

Large-scale multilingual ASR models like Whisper excel in high-resource settings but face challenges in low-resource scenarios, such as rare languages and code-switching (CS), due to computational costs and catastrophic forgetting. We explore Soft Prompt Tuning (SPT), a parameter-efficient method to enhance CS ASR while preserving prior knowledge. We evaluate two strategies: (1) full fine-tuning (FFT) of both soft prompts and the entire Whisper model, demonstrating improved cross-lingual capabilities compared to traditional methods, and (2) adhering to SPT's original design by freezing model parameters and only training soft prompts. Additionally, we introduce SPT4ASR, a combination of different SPT variants. Experiments on the SEAME and ASRU2019 datasets show that deep prompt tuning is the most effective SPT approach, and our SPT4ASR methods achieve further error reductions in CS ASR, maintaining parameter efficiency similar to LoRA, without degrading performance on existing languages.

语音识别混语处理参数高效提示调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。