一个语音编码器同时支持实时与非实时场景,性能更优。
DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation
- 通过ASR感知的蒸馏机制,实现单模型双模式运行。
- 200M参数模型在流式与非流式任务上分别提升12%和11.6%。
- 适用于需要兼顾实时性与精度的语音系统研发者。
近期语音编码器的发展得益于其与大语言模型结合在多种语音任务中的应用。尽管多数研究聚焦于因果或全上下文编码器,但对同时支持流式与非流式应用场景且保持顶尖性能的探索仍有限。我们提出DuRep,一种双模式语音表征学习框架,使单一语音编码器在无需额外参数或模式特定调整的情况下,高效服务于离线与在线任务。杜雷普-200M(200M参数)在多语言自动语音识别任务中,相比基线编码器,流式与非流式模式分别提升12%与11.6%。将该方法扩展至20亿参数规模的DuRep-2B,在自动语音识别及非语音识别任务上均刷新了性能基准。分析揭示了编码器各层在声学与语义信息间的有趣权衡。
原文摘要 · Abstract (English)
Recent advancements in speech encoders have drawn attention due to their integration with Large Language Models for various speech tasks. While most research has focused on either causal or full-context speech encoders, there's limited exploration to effectively handle both streaming and non-streaming applications, while achieving state-of-the-art performance. We introduce DuRep, a Dual-mode Speech Representation learning setup, which enables a single speech encoder to function efficiently in both offline and online modes without additional parameters or mode-specific adjustments, across downstream tasks. DuRep-200M, our 200M parameter dual-mode encoder, achieves 12% and 11.6% improvements in streaming and non-streaming modes, over baseline encoders on Multilingual ASR. Scaling this approach to 2B parameters, DuRep-2B sets new performance benchmarks across ASR and non-ASR tasks. Our analysis reveals interesting trade-offs between acoustic and semantic information across encoder layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。