arXiv:2412.06967cs.CLcs.SD2024-12中稿 · as SLT 2024 procee…被引 9

用软提示微调提升大模型语音识别的领域适应能力

Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning

  • 分两步软提示微调,增强领域文本适配性
  • 相比基线,词错误率降低最多9%,实体错误率降18%
  • 适合需要高精度领域语音识别的场景

大型语言模型(LLM)的兴起重塑了自动语音识别(ASR)。通过将音频嵌入输入LLM以生成转录,已成为新的最优方案。尽管LLM已基于海量文本训练,高质量领域特定文本仍能显著提升领域适配性能。虽然可通过微调LLM解码器融入更多文本语料,但仅使用无配对提示的纯文本数据微调可能削弱领域知识效果。为此,我们提出一种两阶段软提示微调策略,有效增强领域文本适配。实验表明,该方法在目标领域上实现最高9%的词错误率(WER)相对降低和最高18%的实体错误率(EER)相对降低;结合领域专用语言模型(LM)融合可进一步相对提升EER 2%-5%。

原文摘要 · Abstract (English)

The advent of Large Language Models (LLM) has reformed the Automatic Speech Recognition (ASR). Prompting LLM with audio embeddings to generate transcriptions becomes the new state-of-the-art ASR. Despite LLMs being trained with an extensive amount of text corpora, high-quality domain-specific text data can still significantly enhance ASR performance on domain adaptation tasks. Although LLM-based ASR can naturally incorporate more text corpora by fine-tuning the LLM decoder, fine-tuning such ASR on text-only data without paired prompts may diminish the effectiveness of domain-specific knowledge. To mitigate this issue, we propose a two-step soft prompt fine-tuning strategy that enhances domain-specific text adaptation. Experimental results show that text adaptation with our proposed method achieved a relative up to 9% Word Error Rate (WER) reduction and up to 18% Entity Error Rate (EER) reduction on the target domain compared to the baseline ASR. Combining this with domain-specific Language Model (LM) fusion can further improve the EER by a relative 2-5%

语音识别大模型软提示领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。