arXiv:2508.10456eess.AS2025-08被引 1

探索跨话语语音上下文,提升语音识别模型精度

Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems

  • 提出四种跨话语上下文建模方法,包括特征拼接与嵌入池化
  • 在四组数据上实现最高0.98%的字符错误率降低
  • 适合语音识别、老年语言分析等需要上下文理解的场景

本文研究了四种针对流式与非流式Conformer-Transformer(C-T)语音识别系统的跨话语语音上下文建模方法:i)输入音频特征拼接;ii)跨话语编码器嵌入拼接;iii)跨话语编码器嵌入池化投影;iv)首次应用于C-T模型的新型分块方法。为高效训练,提出一种批处理方案,通过拼接每批次内的语音语句,在保持跨话语顺序的同时减少同步开销。实验在四个基准语音数据集上进行,涵盖三种语言:英语GigaSpeech和中文Wenetspeech用于预训练;英语DementiaBank Pitt和粤语JCCOCC MoCA老年人语音数据集用于领域微调。最佳模型在四个任务中均显著优于无跨话语上下文的基线,绝对词错误率(WER)或字符错误率(CER)分别降低0.9%、1.1%、0.51%和0.98%(相对降低6.0%、5.4%、2.0%和3.4%)。其性能媲美Wav2vec2.0-Conformer、XLSR-128和Whisper模型,表明将跨话语上下文融入当前语音基础模型具有重要潜力。

原文摘要 · Abstract (English)

This paper investigates four types of cross-utterance speech contexts modeling approaches for streaming and non-streaming Conformer-Transformer (C-T) ASR systems: i) input audio feature concatenation; ii) cross-utterance Encoder embedding concatenation; iii) cross-utterance Encoder embedding pooling projection; or iv) a novel chunk-based approach applied to C-T models for the first time. An efficient batch-training scheme is proposed for contextual C-Ts that uses spliced speech utterances within each minibatch to minimize the synchronization overhead while preserving the sequential order of cross-utterance speech contexts. Experiments are conducted on four benchmark speech datasets across three languages: the English GigaSpeech and Mandarin Wenetspeech corpora used in contextual C-T models pre-training; and the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets used in domain fine-tuning. The best performing contextual C-T systems consistently outperform their respective baselines using no cross-utterance speech contexts in pre-training and fine-tuning stages with statistically significant average word error rate (WER) or character error rate (CER) reductions up to 0.9%, 1.1%, 0.51%, and 0.98% absolute (6.0%, 5.4%, 2.0%, and 3.4% relative) on the four tasks respectively. Their performance competitiveness against Wav2vec2.0-Conformer, XLSR-128, and Whisper models highlights the potential benefit of incorporating cross-utterance speech contexts into current speech foundation models.

语音识别上下文建模Conformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。