动态调整语音分块,提升藏语流式识别准确率与速度
Context-Aware Dynamic Chunking for Streaming Tibetan Speech Recognition
- 根据编码状态自适应调整分块大小,灵活捕捉上下文
- 藏语测试集上词错误率降至6.23%,较固定分块提升48.15%
- 适合低延迟藏语语音识别系统开发,尤其关注语言特性
本文提出一种面向安多藏语的流式语音识别框架,基于混合CTC/注意力架构,引入上下文感知的动态分块机制。该策略根据编码状态自适应调整分块宽度,实现灵活感受野、跨分块信息交互,并有效应对不同语速变化,缓解固定分块方法的上下文截断问题。为更好地建模藏语语言特征,构建了基于其正字法原则的词典,提供语言学驱动的建模单元。解码阶段引入外部语言模型,增强语义一致性并提升长句识别效果。实验表明,该框架在测试集上词错误率(WER)达6.23%,相较固定分块基线相对降低48.15%,显著降低识别延迟,且性能接近全局解码。
原文摘要 · Abstract (English)
In this work, we propose a streaming speech recognition framework for Amdo Tibetan, built upon a hybrid CTC/Atten-tion architecture with a context-aware dynamic chunking mechanism. The proposed strategy adaptively adjusts chunk widths based on encoding states, enabling flexible receptive fields, cross-chunk information exchange, and robust adaptation to varying speaking rates, thereby alleviating the context truncation problem of fixed-chunk methods. To further capture the linguistic characteristics of Tibetan, we construct a lexicon grounded in its orthographic principles, providing linguistically motivated modeling units. During decoding, an external language model is integrated to enhance semantic consistency and improve recognition of long sentences. Experimental results show that the proposed framework achieves a word error rate (WER) of 6.23% on the test set, yielding a 48.15% relative improvement over the fixed-chunk baseline, while significantly reducing recognition latency and maintaining performance close to global decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。