arXiv:2605.14340cs.SD2026-05中稿 · Interspeech 2026

用语音文本对齐优化伪音频提示,提升纯文本域适应效果

Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR

论文配图:Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR
图 1 · 摘自论文原文
  • 通过建模语音与文本对齐关系生成更丰富的伪音频提示
  • 在数据稀缺场景下显著降低错误率并提升词汇覆盖率
  • 适合需要低资源语音识别域适应的研究者

基于大语言模型(LLM)的自动语音识别系统通过连接音频编码器与LLM展现强大性能。然而,配对语音与转录数据的匮乏常阻碍其在新领域的适配,使得纯文本域适应至关重要。现有方法通常仅微调LLM或使用伪音频提示,前者忽略声学上下文,后者在数据稀缺时可扩展性差,或仅依赖文本特征导致提示表达力不足。为此,我们提出一种新框架,显式建模语音-文本对齐。该方法高效生成高度表达性的伪音频提示,弥合模态差距,实现有效目标域适配。实验表明,本方法优于现有纯文本方法,在整体错误率和未登录词覆盖方面均有提升。

原文摘要 · Abstract (English)

LLM-based automatic speech recognition models demonstrate strong performance by connecting audio encoders and LLMs. However, data scarcity of paired speech and transcription often hinders their adaptation to new domains, making text-only domain adaptation crucial. Existing methods typically rely on either fine-tuning the LLM alone or employing pseudo-audio prompts. The former neglects essential acoustic context, while the latter either suffers from limited scalability in data-scarce conditions, or yields inexpressive prompts by leveraging only textual features, ignoring audio modality. To address this, we propose an enhanced framework that explicitly models speech-text alignment. Our method efficiently generates highly expressive pseudo-audio prompts that bridges the modality gap, enabling effective target-domain adaptation. Experiments demonstrate that our approach outperforms existing text-only methods, improving both overall error rates and out-of-vocabulary coverage.

语音识别域适应伪提示多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。