融合音素、音节、词片信息,提升孟加拉语语音识别准确率。
Multi-Level Embedding Conformer Framework for Bengali Automatic Speech Recognition
- 多粒度嵌入融合:音素、音节、词片特征联合建模。
- 在低资源下实现WER 10.01%、CER 5.03%的先进性能。
- 适合低资源语言语音识别研究者参考应用。
孟加拉语是超过3亿人使用的形态丰富的低资源语言,给自动语音识别(ASR)带来挑战。本文提出一种基于Conformer-CTC骨干网络的端到端孟加拉语ASR框架,引入多层级嵌入融合机制,融合音素、音节和词片表示。通过将这些语言学嵌入融入声学特征,模型能捕捉细微的语音线索与高层上下文模式。架构采用早期和晚期Conformer阶段,预处理包括静音剪裁、重采样、对数梅尔谱图提取及SpecAugment增强。实验结果表明,该模型表现优异,达到10.01%的词错误率(WER)和5.03%的字符错误率(CER),验证了多粒度语言信息与声学建模结合的有效性,为低资源ASR开发提供可扩展方案。
原文摘要 · Abstract (English)
Bengali, spoken by over 300 million people, is a morphologically rich and lowresource language, posing challenges for automatic speech recognition (ASR). This research presents an end-to-end framework for Bengali ASR, building on a Conformer-CTC backbone with a multi-level embedding fusion mechanism that incorporates phoneme, syllable, and wordpiece representations. By enriching acoustic features with these linguistic embeddings, the model captures fine-grained phonetic cues and higher-level contextual patterns. The architecture employs early and late Conformer stages, with preprocessing steps including silence trimming, resampling, Log-Mel spectrogram extraction, and SpecAugment augmentation. The experimental results demonstrate the strong potential of the model, achieving a word error rate (WER) of 10.01% and a character error rate (CER) of 5.03%. These results demonstrate the effectiveness of combining multi-granular linguistic information with acoustic modeling, providing a scalable approach for low-resource ASR development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。