arXiv:2410.10219cs.CL2024-10

用音译法解决濒危语言翻译难题,首次实现查克马语机器翻译。

ChakmaNMT: Machine Translation for a Low-Resource and Endangered Language via Transliteration

  • 通过音译桥接查克马语与孟加拉语,缓解数据稀缺问题。
  • 音译模型使翻译性能显著提升,跨方向差异明显。
  • 适合语言保护研究者和低资源语言技术开发者。

我们首次系统研究了查克马语——一种濒危且极度低资源的印度-雅利安语系语言的机器翻译,旨在支持语言可及性与保护。提出新的查克马-孟加拉语平行语料库与单语语料库,并构建三语(查克马-孟加拉语-英语)评估基准。针对书写系统差异与数据匮乏,提出字符级音译框架,利用查克马语与孟加拉语在拼写与语音上的紧密关联,在保留语义的同时实现从孟加拉语及多语言预训练模型的有效迁移。对比从头训练、微调预训练模型及基于上下文学习的大语言模型,结果表明音译至关重要,微调与上下文学习显著优于从头训练,且翻译方向存在明显不对称性。

原文摘要 · Abstract (English)

We present the first systematic study of machine translation for Chakma, an endangered and extremely low-resource Indo-Aryan language, with the goal of supporting language access and preservation. We introduce a new Chakma-Bangla parallel and monolingual dataset, along with a trilingual Chakma-Bangla-English benchmark for evaluation. To address script mismatch and data scarcity, we propose a character-level transliteration framework that exploits the close orthographic and phonological relationship between Chakma and Bangla, preserving semantic content while enabling effective transfer from Bangla and multilingual pretrained models. We benchmark from-scratch MT, fine-tuned pretrained models, and large language models via in-context learning. Results show that transliteration is essential and that fine-tuning and in-context learning substantially outperform from-scratch baselines, with strong asymmetry across translation directions.

机器翻译濒危语言音译低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。