arXiv:2601.01778cs.CL2026-01中稿 · LoResLM workshop, …被引 4

解决孟加拉语语音转写中的方言与数字难题

BanglaIPA: Towards Robust Text-to-IPA Transcription with Contextual Rewriting in Bengali

  • 用字符级词典+词级对齐,提升转写鲁棒性
  • 在6种方言上误差率仅11.4%,性能提升58.4%-78.7%
  • 适合需要处理方言和数字的孟加拉语语音系统开发者

尽管孟加拉语使用广泛,但缺乏能有效处理标准语和地方方言的自动国际音标(IPA)转写系统。现有方法难以应对方言差异、数字表达,且对未见词汇泛化能力差。为此,我们提出BanglaIPA,一种结合字符级词汇表与词级对齐的新型IPA生成系统。该系统能准确处理孟加拉语数字,并在标准语及六种地区变体上表现优异。通过预计算词到IPA映射字典,显著提升推理效率。在DUAL-IPA数据集的标准孟加拉语及六种方言上评估,结果表明,BanglaIPA相比基线模型性能提升58.4%-78.7%,整体平均词错误率为11.4%,展现出强大的语音转写鲁棒性。

原文摘要 · Abstract (English)

Despite its widespread use, Bengali lacks a robust automated International Phonetic Alphabet (IPA) transcription system that effectively supports both standard language and regional dialectal texts. Existing approaches struggle to handle regional variations, numerical expressions, and generalize poorly to previously unseen words. To address these limitations, we propose BanglaIPA, a novel IPA generation system that integrates a character-based vocabulary with word-level alignment. The proposed system accurately handles Bengali numerals and demonstrates strong performance across regional dialects. BanglaIPA improves inference efficiency by leveraging a precomputed word-to-IPA mapping dictionary for previously observed words. The system is evaluated on the standard Bengali and six regional variations of the DUAL-IPA dataset. Experimental results show that BanglaIPA outperforms baseline IPA transcription models by 58.4-78.7% and achieves an overall mean word error rate of 11.4%, highlighting its robustness in phonetic transcription generation for the Bengali language.

语音转写孟加拉语方言处理IPA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。