arXiv:2601.13044cs.CL2026-01被引 2

轻量级模型实现低延迟泰语语音识别,效果媲美大模型。

Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition

  • 采用FastConformer-Transducer架构,仅115M参数实现低延迟。
  • 计算成本降低45倍,准确率接近Whisper Large-v3。
  • 专为泰语设计文本规范化与方言适配,适合实时应用研究者。

大型编码器-解码器模型如Whisper在离线转录中表现优异,但在流式应用中因高延迟难以实用。由于预训练模型的可得性,当前开放泰语语音识别领域仍以这些离线架构为主,缺乏高效流式解决方案。我们提出Typhoon ASR Real-time,一个115M参数的FastConformer-Transducer模型,实现低延迟泰语语音识别。实验表明,严格的文本规范化可达到与模型扩容相当的效果:该紧凑模型相比Whisper Large-v3降低45倍计算成本,同时保持相近准确率。我们的规范化流程解决了泰语转录中的系统性歧义问题,包括上下文相关的数字表达和重复标记(mai yamok),生成一致的训练目标。此外,我们引入两阶段课程学习方法,实现伊桑语(东北方言)适应,同时保留中部泰语性能。为解决泰语语音识别的可复现性挑战,我们发布Typhoon ASR Benchmark,包含符合泰语语言学规范的人工标注数据集及标准化评估协议,为研究社区提供基准支持。

原文摘要 · Abstract (English)

Large encoder-decoder models like Whisper achieve strong offline transcription but remain impractical for streaming applications due to high latency. However, due to the accessibility of pre-trained checkpoints, the open Thai ASR landscape remains dominated by these offline architectures, leaving a critical gap in efficient streaming solutions. We present Typhoon ASR Real-time, a 115M-parameter FastConformer-Transducer model for low-latency Thai speech recognition. We demonstrate that rigorous text normalization can match the impact of model scaling: our compact model achieves a 45x reduction in computational cost compared to Whisper Large-v3 while delivering comparable accuracy. Our normalization pipeline resolves systemic ambiguities in Thai transcription --including context-dependent number verbalization and repetition markers (mai yamok) --creating consistent training targets. We further introduce a two-stage curriculum learning approach for Isan (north-eastern) dialect adaptation that preserves Central Thai performance. To address reproducibility challenges in Thai ASR, we release the Typhoon ASR Benchmark, a gold-standard human-labeled datasets with transcriptions following established Thai linguistic conventions, providing standardized evaluation protocols for the research community.

语音识别低延迟泰语流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。