用自监督预训练+大模型提升中文方言语音识别效果
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
- 在30万小时无标签方言数据上预训练Data2vec2模型
- 在4万小时有标签数据上对齐训练,方言识别达最优性能
- 适用于低资源方言语音识别,适合语音技术研究者
大规模训练语料显著提升了语音识别(ASR)模型性能。然而,由于数据稀缺,中文口音和方言仍是多数ASR模型的挑战。自监督学习的进步表明,在低资源场景下,结合自监督预训练与大语言模型(LLM)可有效提升ASR表现。本文旨在探究该范式在中文方言中的有效性:我们在30万小时未标注的方言及口音语音数据上预训练Data2vec2模型,并在4万小时标注数据集上进行对齐训练。系统评估了不同投影层与LLM组合在普通话、方言及口音语音识别中的影响。实验结果在多个方言数据集(包括Kespeech)上达到当前最优水平。相关工作将开源,以促进可复现研究。
原文摘要 · Abstract (English)
Large-scale training corpora have significantly improved the performance of ASR models. Unfortunately, due to the relative scarcity of data, Chinese accents and dialects remain a challenge for most ASR models. Recent advancements in self-supervised learning have shown that self-supervised pre-training, combined with large language models (LLM), can effectively enhance ASR performance in low-resource scenarios. We aim to investigate the effectiveness of this paradigm for Chinese dialects. Specifically, we pre-train a Data2vec2 model on 300,000 hours of unlabeled dialect and accented speech data and do alignment training on a supervised dataset of 40,000 hours. Then, we systematically examine the impact of various projectors and LLMs on Mandarin, dialect, and accented speech recognition performance under this paradigm. Our method achieved SOTA results on multiple dialect datasets, including Kespeech. We will open-source our work to promote reproducible research
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。