用迁移学习解决少数民族语言翻译难题,促进越南两族文化互通
Towards Cultural Bridge by Bahnaric-Vietnamese Translation Using Transfer Learning of Sequence-To-Sequence Pre-training Language Model
- 基于序列到序列预训练模型,利用越语资源迁移至巴拿语翻译
- 在少量双语数据下实现有效翻译,提升资源匮乏语言的可用性
- 适合关注小语种、文化保护与跨语言技术的研究者
本文致力于实现巴拿语与越语之间的翻译,以促进越南两个民族间的文化交流。然而,巴拿语作为原始语料稀缺的语言,面临词汇、语法、对话模式及双语语料库严重不足的问题,阻碍了训练数据的收集。为此,我们采用序列到序列预训练语言模型的迁移学习方法:首先利用预训练越语模型捕捉其语言特征;特别地,为支持机器翻译,选用序列到序列架构而非仅编码器(如BERT)或仅解码器(如GPT)模型。鉴于两语言间显著相似性,我们使用现有有限的越-巴拿语双语文本继续训练模型,实现从语言模型到翻译任务的迁移。该方法缓解了双语资源不平衡问题,并优化了训练与计算效率。此外,通过数据增强生成额外语料,并引入启发式方法提升翻译精度。实验验证该方法在巴拿语-越语翻译中表现高效,有助于语言扩展与保护,推动两族间更深入的理解与交流。
原文摘要 · Abstract (English)
This work explores the journey towards achieving Bahnaric-Vietnamese translation for the sake of culturally bridging the two ethnic groups in Vietnam. However, translating from Bahnaric to Vietnamese also encounters some difficulties. The most prominent challenge is the lack of available original Bahnaric resources source language, including vocabulary, grammar, dialogue patterns and bilingual corpus, which hinders the data collection process for training. To address this, we leverage a transfer learning approach using sequence-to-sequence pre-training language model. First of all, we leverage a pre-trained Vietnamese language model to capture the characteristics of this language. Especially, to further serve the purpose of machine translation, we aim for a sequence-to-sequence model, not encoder-only like BERT or decoder-only like GPT. Taking advantage of significant similarity between the two languages, we continue training the model with the currently limited bilingual resources of Vietnamese-Bahnaric text to perform the transfer learning from language model to machine translation. Thus, this approach can help to handle the problem of imbalanced resources between two languages, while also optimizing the training and computational processes. Additionally, we also enhanced the datasets using data augmentation to generate additional resources and defined some heuristic methods to help the translation more precise. Our approach has been validated to be highly effective for the Bahnaric-Vietnamese translation model, contributing to the expansion and preservation of languages, and facilitating better mutual understanding between the two ethnic people.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。