用音标和对齐信息提升缅甸语低资源语音识别纠错效果
ASR Error Correction in Low-Resource Burmese with Alignment-Enhanced Transformers using Phonetic Features
- 融合国际音标与对齐信息的Transformer模型进行纠错
- 纠错后词错误率从51.56降至39.82(增广前)
- 适合低资源语言语音识别优化,尤其关注缅甸语
本文研究序列到序列Transformer模型在低资源缅甸语语音识别纠错中的应用,重点分析了国际音标(IPA)与对齐信息等特征融合策略。据我们所知,这是首个针对缅甸语语音识别纠错的研究。评估五种ASR骨干模型后发现,所提纠错方法在词级和字符级准确率上均优于基线。结合IPA与对齐特征的纠错模型,使未增广情况下的平均词错误率(WER)从51.56降至39.82,增广后从51.56降至43.59;同时chrF++得分由0.5864提升至0.627,显著优于无纠错的基线输出。结果表明,纠错机制有效且特征设计对低资源场景下语音识别性能提升至关重要。
原文摘要 · Abstract (English)
This paper investigates sequence-to-sequence Transformer models for automatic speech recognition (ASR) error correction in low-resource Burmese, focusing on different feature integration strategies including IPA and alignment information. To our knowledge, this is the first study addressing ASR error correction specifically for Burmese. We evaluate five ASR backbones and show that our ASR Error Correction (AEC) approaches consistently improve word- and character-level accuracy over baseline outputs. The proposed AEC model, combining IPA and alignment features, reduced the average WER of ASR models from 51.56 to 39.82 before augmentation (and 51.56 to 43.59 after augmentation) and improving chrF++ scores from 0.5864 to 0.627, demonstrating consistent gains over the baseline ASR outputs without AEC. Our results highlight the robustness of AEC and the importance of feature design for improving ASR outputs in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。