arXiv:2606.21268cs.SD2026-06中稿 · Interspeech 2026

提升语音模型在流式与非流式下的性能一致性

Online Predictive Coding for Dual-Mode Self-Supervised Speech Model

论文配图:Online Predictive Coding for Dual-Mode Self-Supervised Speech Model
图 1 · 摘自论文原文
  • 引入在线预测编码,用多步未来预测优化注册表
  • 在160毫秒延迟下,测试集词错误率降低至3.40%和9.65%
  • 适合需要低延迟语音识别的工业场景

双模式自监督语音模型旨在同时适应流式与非流式输入,但因注意力作用范围不同导致优化困难。此前我们提出在线注册表以补偿流式模式中缺失的未来上下文,但效果有限。本文提出两项改进:(1) 在线预测编码(OPC),通过多步未来预测对注册表进行正则化;(2) 双模式层归一化,增强训练稳定性。在LibriSpeech和WSJ数据集上微调用于语音识别,结果表明,OPC有效缩小了在线与离线性能差距:在160毫秒延迟下,LibriSpeech测试集清洁数据词错误率从3.65%降至3.40%,测试集其他数据从10.15%降至9.65%。

原文摘要 · Abstract (English)

Dual-mode self-supervised speech models are pre-trained to handle streaming and non-streaming conditions simultaneously. However, their attention is computed over different context ranges, which often makes optimization difficult. In previous work, we proposed online registers, additional tokens intended to compensate for missing future context in streaming mode, but the gains remained limited. To address these issues, we introduce two improvements for robust dual-mode pre-training: (1) Online Predictive Coding (OPC), which regularizes the registers through multi-step future prediction, and (2) Dual-mode Layer Normalization, which stabilizes optimization. We fine-tune the proposed dual-mode self-supervised speech models for speech recognition on LibriSpeech and WSJ. Results show that OPC consistently reduces the online-offline performance gap; at 160 ms latency on LibriSpeech, word error rates improve from 3.65% to 3.40% on test-clean and from 10.15% to 9.65% on test-other.

语音识别自监督学习流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。