arXiv:2411.19726cs.CLcs.LG2024-11被引 2

首个桑塔利语到英语的翻译模型,解决低资源语言技术缺失问题。

Towards Santali Linguistic Inclusion: Building the First Santali-to-English Translation Model using mT5 Transformer and Data Augmentation

  • 基于mT5 Transformer与数据增强构建翻译模型
  • 在低资源下实现有效翻译,性能优于桑塔利-孟加拉语对
  • 为濒危语言数字化提供可复用的技术路径

印度、孟加拉、不丹和尼泊尔约有七百万使用者讲桑塔利语,是南亚语系蒙达语支中使用人数第三多的语言。尽管如此,桑塔利语仍缺乏全球认可度,目前尚无相关机器翻译模型。本文旨在推动桑塔利语进入自然语言处理领域,探索基于现有桑塔利语语料库构建翻译模型的可行性。研究成功应对了低资源挑战,在采用mT5 Transformer并结合数据增强的情况下,实现了具有前景的翻译性能。结果表明,相较于未预训练的模型,mT5在桑塔利-英语平行语料上表现更优,验证了迁移学习在该语言上的有效性;同时,因mT5在英文数据上训练更充分,其在桑塔利-英语对上的表现优于桑塔利-孟加拉语对。此外,数据增强显著提升了模型性能。

原文摘要 · Abstract (English)

Around seven million individuals in India, Bangladesh, Bhutan, and Nepal speak Santali, positioning it as nearly the third most commonly used Austroasiatic language. Despite its prominence among the Austroasiatic language family's Munda subfamily, Santali lacks global recognition. Currently, no translation models exist for the Santali language. Our paper aims to include Santali to the NPL spectrum. We aim to examine the feasibility of building Santali translation models based on available Santali corpora. The paper successfully addressed the low-resource problem and, with promising results, examined the possibility of creating a functional Santali machine translation model in a low-resource setup. Our study shows that Santali-English parallel corpus performs better when in transformers like mt5 as opposed to untrained transformers, proving that transfer learning can be a viable technique that works with Santali language. Besides the mT5 transformer, Santali-English performs better than Santali-Bangla parallel corpus as the mT5 has been trained in way more English data than Bangla data. Lastly, our study shows that with data augmentation, our model performs better.

机器翻译低资源语言mT5数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。