用少量电话数据适配语音基础模型,实现多说话人语音识别的高效泛化。
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
- 仅用电话数据微调语音基础模型,实现多说话人识别。
- 参数越少,整体词错误率越低,反直觉但有效。
- 无需额外训练即可在会议数据上表现良好,适合资源受限场景。
语音基础模型已在数百种语言的自动语音识别(ASR)任务中达到顶尖性能。然而,由于数据稀缺与稀疏性,多说话人语音识别仍是挑战。本文提出方法,仅使用电话通话数据即可适配语音基础模型完成多说话人ASR任务。值得注意的是,经适配的模型在未经过任何微调的情况下,也能在会议数据上表现优异,展现出良好的泛化能力。通过多项消融实验分析不同参数与策略对性能的影响,结果表明:更少的参数反而带来更低的整体词错误率(cpWER),这一反直觉现象为在极低标注数据下适配语音基础模型提供了新思路。
原文摘要 · Abstract (English)
Speech foundation models have achieved state-of-the-art (SoTA) performance across various tasks, such as automatic speech recognition (ASR) in hundreds of languages. However, multi-speaker ASR remains a challenging task for these models due to data scarcity and sparsity. In this paper, we present approaches to enable speech foundation models to process and understand multi-speaker speech with limited training data. Specifically, we adapt a speech foundation model for the multi-speaker ASR task using only telephonic data. Remarkably, the adapted model also performs well on meeting data without any fine-tuning, demonstrating the generalization ability of our approach. We conduct several ablation studies to analyze the impact of different parameters and strategies on model performance. Our findings highlight the effectiveness of our methods. Results show that less parameters give better overall cpWER, which, although counter-intuitive, provides insights into adapting speech foundation models for multi-speaker ASR tasks with minimal annotated data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。