用SSL离散令牌实现多语言语音识别,效果媲美传统特征
Exploring SSL Discrete Tokens for Multilingual ASR
- 用多种主流SSL模型生成离散语音令牌,用于多语言语音识别
- 在7种语言上平均词错误率降低0.31%(测试集),波兰语降6.82%
- 适合追求高效多语言语音识别的开发者与研究者
随着自监督学习(SSL)在语音任务中的发展,利用SSL生成的离散令牌进行自动语音识别(ASR)日益受到关注,因其能实现更快速的处理。然而,以往研究主要聚焦于使用Fbank特征的多语言ASR或基于离散令牌的英语ASR,缺乏对离散令牌在多语言场景下的适配探索。本研究全面比较了多种领先SSL模型在多个语言领域生成的离散令牌性能。实验表明,在七个语言领域的单语和多语ASR任务中,离散令牌的表现可媲美基于Fbank特征的系统:在开发集和测试集上分别实现0.31%和1.76%的绝对词错误率(WER)下降(相对下降分别为2.80%和15.70%),尤其在波兰语测试集上绝对下降达6.82%(相对下降41.48%)。
原文摘要 · Abstract (English)
With the advancement of Self-supervised Learning (SSL) in speech-related tasks, there has been growing interest in utilizing discrete tokens generated by SSL for automatic speech recognition (ASR), as they offer faster processing techniques. However, previous studies primarily focused on multilingual ASR with Fbank features or English ASR with discrete tokens, leaving a gap in adapting discrete tokens for multilingual ASR scenarios. This study presents a comprehensive comparison of discrete tokens generated by various leading SSL models across multiple language domains. We aim to explore the performance and efficiency of speech discrete tokens across multiple language domains for both monolingual and multilingual ASR scenarios. Experimental results demonstrate that discrete tokens achieve comparable results against systems trained on Fbank features in ASR tasks across seven language domains with an average word error rate (WER) reduction of 0.31% and 1.76% absolute (2.80% and 15.70% relative) on dev and test sets respectively, with particularly WER reduction of 6.82% absolute (41.48% relative) on the Polish test set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。