通过交错语音文本序列提升ASR性能,让大模型更有效利用文本知识。
Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving
- 构建词级与段级交错的语音-文本对,增强文本知识在语音识别中的利用。
- 38小时数据下实体识别准确率优于纯语音和简单联合训练基线。
- 用真实领域文本即可达媲美合成数据效果,简化领域适配,适合医疗等专业场景。
语音-大语言模型(Speech-LLM)融合借助海量文本预训练展现出潜力,但其对自动语音识别(ASR)的具体价值尚不明确。我们发现,随着监督式ASR训练数据增加,大模型先验的贡献逐渐减弱,而简单的语音-文本联合训练未能充分挖掘文本知识。为此,我们提出面向ASR的交错预训练方法(JSTIP),在对齐的语音-文本对中构建词级与段级交错的连续输入序列,适用于支持连续输入的Speech-LLM架构。在38,000小时的ASR数据上实验表明,相较于仅用语音或简单联合训练的基线,JSTIP实现了稳定的实体识别准确率提升。使用真实领域转录文本时,其表现与合成语音-文本对相当,显著简化了领域迁移。得益于文本预训练与领域文本数据,JSTIP在医疗实体识别任务上已具备与开源ASR及Speech-LLM系统竞争的实力。零样本语音问答行为进一步表明,交错设计缩小了语音与文本模态差距,并保留了大模型生成先验,这可能是其在ASR任务中提升实体识别的关键原因。
原文摘要 · Abstract (English)
Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data increases, the contribution of LLM priors becomes less evident, and simple speech-text joint training under-utilizes textual knowledge. We therefore propose Joint Speech-Text Interleaved Pretraining (JSTIP), an ASR-oriented pretraining strategy that constructs word-level and segment-level interleaved speech-text sequences within aligned pairs for speech-LLM architectures that accept continuous inputs. Experiments on 38k hours of ASR data show consistent entity accuracy improvement compared to ASR-only and joint speech-text training baselines. JSTIP achieves on-par entity recognition performance using domain transcription text compared to synthetic speech-text pairs, simplifying domain adaptation. Benefiting from textual pretraining and domain text data, JSTIP is competitive with open-source ASR and Speech-LLM systems in medical entity recognition. The zero-shot speech question answering behaviors further suggest that interleaving reduces the speech-text modality gap and preserves the LLM generative prior, which is likely the reason for the entity improvements on the ASR task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。