用少量语音数据显著提升大模型语音识别的跨域适应能力
Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR
- 引入混合批量训练,融合少量语音与大量文本数据
- 仅用10%目标域语音(不足4小时)即达全量数据效果
- 适合资源受限下提升语音识别泛化性能的研究者
传统端到端语音识别系统依赖成对的语音-文本数据进行领域适配。近期基于大语言模型的语音识别架构通过投影模块将语音编码器与大语言模型连接,实现仅用文本数据的适配。但该方法引入模态差距,因大语言模型未接触语音投影器产生的噪声表示。本文研究少量语音是否可缓解此差异。比较三种策略:纯文本适配、成对语音-文本适配、混合批量(MB)训练。在同域与跨域设置下实验表明,即使少量语音也能持续提升性能。特别地,仅使用10%目标域语音(少于4小时)的混合批量方法,达到或优于使用完整数据集的传统语音识别微调效果,表明少量语音能提供强模态对齐信号。
原文摘要 · Abstract (English)
Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection module, enabling adaptation with text-only data. However, this introduces a modality gap, as the LLM is not exposed to the noisy representations produced by the speech projector. We investigate whether small amounts of speech can mitigate this mismatch. We compare three strategies: text-only adaptation, paired speech-text adaptation, and mixed batching (MB), which combines both. Experiments in in-domain and out-of-domain settings show that even limited speech consistently improves performance. Notably, MB using only 10% of the target-domain (less than 4 hours) speech achieves word error rates comparable to, or better than, conventional ASR fine-tuning with the full dataset, indicating that small amounts of speech provide a strong modality-alignment signal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。