arXiv:2506.21577cs.CLcs.AI2025-06中稿 · Interspeech 2025被引 4

通过智能提示调优,让多语种语音识别模型更高效地新增语言。

Language-Aware Prompt Tuning for Parameter-Efficient Seamless Language Expansion in Multilingual ASR

论文配图:Language-Aware Prompt Tuning for Parameter-Efficient Seamless Language Expansion in Multilingual ASR
图 1 · 摘自论文原文
  • 在编码器和解码器都加软提示,增强特征提取与解码能力。
  • 利用语言相似性区分共通与特有特征,提升新语言识别准确率16%。
  • 轻量级设计适合持续学习,适合需要动态扩展语言的系统使用。

近年来,以Whisper为代表的端到端大规模多语种自动语音识别(ASR)模型推动了该领域的发展。然而,语言干扰问题以及在不降低性能的前提下扩展至未见语言(语言扩展)的挑战依然存在。本文提出三项贡献:1)完整软提示调优(Entire SPT),将软提示应用于编码器和解码器,提升特征提取与解码能力;2)语言感知提示调优(LAPT),利用跨语言相似性,通过轻量级提示矩阵编码共享与语言特定特征;3)SPT-Whisper工具包,将SPT集成至Whisper,支持高效连续学习。在FLEURS三个语言上的实验表明,Entire SPT与LAPT在语言扩展任务中分别比仅解码器调优(Decoder SPT)提升5.0%和16.0%,为多语种动态语音识别模型提供低计算开销的高效解决方案。

原文摘要 · Abstract (English)

Recent advancements in multilingual automatic speech recognition (ASR) have been driven by large-scale end-to-end models like Whisper. However, challenges such as language interference and expanding to unseen languages (language expansion) without degrading performance persist. This paper addresses these with three contributions: 1) Entire Soft Prompt Tuning (Entire SPT), which applies soft prompts to both the encoder and decoder, enhancing feature extraction and decoding; 2) Language-Aware Prompt Tuning (LAPT), which leverages cross-lingual similarities to encode shared and language-specific features using lightweight prompt matrices; 3) SPT-Whisper, a toolkit that integrates SPT into Whisper and enables efficient continual learning. Experiments across three languages from FLEURS demonstrate that Entire SPT and LAPT outperform Decoder SPT by 5.0% and 16.0% in language expansion tasks, respectively, providing an efficient solution for dynamic, multilingual ASR models with minimal computational overhead.

语音识别提示调优多语种轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。