用大模型提升泰语语音识别,解决数据少、算力高的难题
Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR
- 用自进化标签优化策略提升弱标注数据质量
- 序列压缩模块降低计算量,保持识别准确率
- 首个基于大模型的泰语语音识别系统,适合低资源场景
尽管取得显著进展,低资源场景下的语音识别仍面临高质量数据稀缺和高计算需求两大挑战。本文提出EThai-ASR,首个将大语言模型(LLM)应用于泰语语音识别的高效系统。该系统包含语音编码器、连接模块和泰语大模型解码器。为缓解数据稀缺并获得强大语音编码器,EThai-ASR引入自进化数据精炼策略,优化弱标签,提升编码器性能。此外,我们在连接模块中设计可插拔的序列压缩模块,提供三种模式以减少序列长度,从而降低计算开销,同时保持良好识别效果。大量实验表明,EThai-ASR在多个数据集上达到当前最优准确率。我们公开了经过精炼的文本转录结果,以推动后续研究。
原文摘要 · Abstract (English)
Despite remarkable achievements, automatic speech recognition (ASR) in low-resource scenarios still faces two challenges: high-quality data scarcity and high computational demands. This paper proposes EThai-ASR, the first to apply large language models (LLMs) to Thai ASR and create an efficient LLM-based ASR system. EThai-ASR comprises a speech encoder, a connection module and a Thai LLM decoder. To address the data scarcity and obtain a powerful speech encoder, EThai-ASR introduces a self-evolving data refinement strategy to refine weak labels, yielding an enhanced speech encoder. Moreover, we propose a pluggable sequence compression module used in the connection module with three modes designed to reduce the sequence length, thus decreasing computational demands while maintaining decent performance. Extensive experiments demonstrate that EThai-ASR has achieved state-of-the-art accuracy in multiple datasets. We release our refined text transcripts to promote further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。