用50小时标注数据+7万小时无标签数据,训练出性能超越商用系统的越南语语音识别模型。
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
- 通过多轮自监督学习,利用7万小时无标签数据预训练,仅需50小时标注数据微调。
- 在真实数据上性能超越Whisper Large-v3和商用系统,且模型轻量、成本低。
- 适合研究低资源语言语音识别的学者与开发者,开源代码助力社区发展。
自动语音识别(ASR)虽取得显著进展,但严重依赖大规模标注数据,而越南语等低资源语言数据稀缺。现有系统如Whisper、USM和MMS虽表现良好,但在训练成本、延迟和可及性方面仍不足。为此,我们提出VietASR,一种新颖的ASR训练流程,充分利用大量无标签数据与少量标注数据。通过在大规模无标签数据集上进行多轮ASR偏置的自监督学习,VietASR提供了一种低成本、实用的提升ASR性能的方法。实验表明,使用70,000小时无标签数据预训练并仅用50小时标注数据微调,即可获得轻量但强大的ASR模型,在真实数据上优于Whisper Large-v3和商业ASR系统。代码与模型将开源,以促进低资源语言ASR研究。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) has made remarkable progress but heavily relies on large-scale labeled data, which is scarce for low-resource languages like Vietnamese. While existing systems such as Whisper, USM, and MMS achieve promising performance, their efficacy remains inadequate in terms of training costs, latency, and accessibility. To address these issues, we propose VietASR, a novel ASR training pipeline that leverages vast amounts of unlabeled data and a small set of labeled data. Through multi-iteration ASR-biased self-supervised learning on a large-scale unlabeled dataset, VietASR offers a cost-effective and practical solution for enhancing ASR performance. Experiments demonstrate that pre-training on 70,000-hour unlabeled data and fine-tuning on merely 50-hour labeled data yield a lightweight but powerful ASR model. It outperforms Whisper Large-v3 and commercial ASR systems on real-world data. Our code and models will be open-sourced to facilitate research in low-resource ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。