arXiv:2501.05749cs.CL2025-01中稿 · 2024 27th Internat…被引 7

用神经模型将标准孟加拉语转为地方方言,提升技术对语言多样性的支持。

Bridging Dialects: Translating Standard Bangla to Regional Variants Using Neural Models

  • 用NMT模型将标准孟加拉语转为5个地区方言,基于自建数据集
  • BanglaT5表现最优,字符错误率12.3%,词错误率15.7%
  • 适合关注本土语言技术、文化保护的研究者与开发者

孟加拉语包含众多地区方言,丰富了其文化内涵。由于词汇、发音和句法在吉大港、锡尔赫特、巴里沙尔、诺阿哈尔和米门辛格等地区差异显著,标准孟加拉语向方言的翻译面临挑战。这些方言虽对本地身份至关重要,但在技术应用中却缺乏体现。本研究通过神经机器翻译(NMT)模型(包括BanglaT5、mT5和mBART50)解决这一缺口,将标准孟加拉语翻译为上述方言。研究基于包含32,500个句子的“Vashantor”数据集进行微调,并使用字符错误率(CER)和词错误率(WER)评估性能。结果表明,BanglaT5表现最佳,CER为12.3%,WER为15.7%,有效捕捉方言特征。该研究推动了包容性语言技术的发展,促进语言多样性保护。

原文摘要 · Abstract (English)

The Bangla language includes many regional dialects, adding to its cultural richness. The translation of Bangla Language into regional dialects presents a challenge due to significant variations in vocabulary, pronunciation, and sentence structure across regions like Chittagong, Sylhet, Barishal, Noakhali, and Mymensingh. These dialects, though vital to local identities, lack of representation in technological applications. This study addresses this gap by translating standard Bangla into these dialects using neural machine translation (NMT) models, including BanglaT5, mT5, and mBART50. The work is motivated by the need to preserve linguistic diversity and improve communication among dialect speakers. The models were fine-tuned using the "Vashantor" dataset, containing 32,500 sentences across various dialects, and evaluated through Character Error Rate (CER) and Word Error Rate (WER) metrics. BanglaT5 demonstrated superior performance with a CER of 12.3% and WER of 15.7%, highlighting its effectiveness in capturing dialectal nuances. The outcomes of this research contribute to the development of inclusive language technologies that support regional dialects and promote linguistic diversity.

语言翻译方言生成NMT孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。