构建非正式语料库,提升孟加拉语机器翻译能力
Advancing Bangla Machine Translation Through Informal Datasets
- 从社交媒体等非正式来源构建孟加拉语-英语语料库
- 首次系统性针对非正式孟加拉语进行机器翻译优化
- 适合关注低资源语言和数字包容性的研究者
孟加拉语是全球第六大使用语言,母语者约2.34亿人。然而,开源孟加拉语机器翻译进展有限,多数网络资源为英文且未翻译成孟加拉语,使数百万用户无法获取关键信息。现有研究多聚焦正式语言,忽视日常使用的非正式语言,主因是缺乏成对的孟加拉语-英语数据及先进翻译模型。本文探索当前最先进的模型,通过从社交媒体和对话文本中构建非正式语料库,推动孟加拉语机器翻译发展,重点提升非正式语言翻译性能,增强孟加拉语使用者在数字世界的资讯可及性。
原文摘要 · Abstract (English)
Bangla is the sixth most widely spoken language globally, with approximately 234 million native speakers. However, progress in open-source Bangla machine translation remains limited. Most online resources are in English and often remain untranslated into Bangla, excluding millions from accessing essential information. Existing research in Bangla translation primarily focuses on formal language, neglecting the more commonly used informal language. This is largely due to the lack of pairwise Bangla-English data and advanced translation models. If datasets and models can be enhanced to better handle natural, informal Bangla, millions of people will benefit from improved online information access. In this research, we explore current state-of-the-art models and propose improvements to Bangla translation by developing a dataset from informal sources like social media and conversational texts. This work aims to advance Bangla machine translation by focusing on informal language translation and improving accessibility for Bangla speakers in the digital world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。