arXiv:2506.05671eess.AScs.CL2025-06中稿 · publication in ASR…被引 15

仅用文本就能让语音大模型快速适配新领域,无需音频数据。

Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning

  • 用目标域文本微调,不依赖配对语音数据。
  • 在多个数据集上表现接近全量语音微调,性能下降小。
  • 适合语音数据稀缺但文本易得的场景,避免遗忘旧领域知识。

近期自动语音识别(ASR)发展将语音编码器与大语言模型(LLMs)通过投影结合,形成性能强劲的语音大模型(Speech LLM)。然而,在语音-文本配对数据稀少的低资源环境下,将其适配至新领域仍具挑战。本文提出一种仅使用未配对目标域文本的文本微调策略,无需额外音频输入。为保持语音-文本对齐,引入实时评估机制以指导微调过程。实验在LibriSpeech、SlideSpeech和Medical数据集上验证,该方法在保持源域性能的同时实现有效领域迁移,识别性能接近全量语音-文本微调,且无显著性能退化,展现出文本仅微调在低资源语音领域适应中的潜力。

原文摘要 · Abstract (English)

Recent advances in automatic speech recognition (ASR) have combined speech encoders with large language models (LLMs) through projection, forming Speech LLMs with strong performance. However, adapting them to new domains remains challenging, especially in low-resource settings where paired speech-text data is scarce. We propose a text-only fine-tuning strategy for Speech LLMs using unpaired target-domain text without requiring additional audio. To preserve speech-text alignment, we introduce a real-time evaluation mechanism during fine-tuning. This enables effective domain adaptation while maintaining source-domain performance. Experiments on LibriSpeech, SlideSpeech, and Medical datasets show that our method achieves competitive recognition performance, with minimal degradation compared to full audio-text fine-tuning. It also improves generalization to new domains without catastrophic forgetting, highlighting the potential of text-only fine-tuning for low-resource domain adaptation of ASR.

语音大模型低资源文本微调领域自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。