arXiv:2509.20209cs.CLcs.AI2025-09

用自定义分词和多语言模型提升提格雷尼亚语翻译质量

Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks

  • 结合语言特异性分词与迁移学习提升低资源语言翻译
  • 自定义分词使翻译性能显著优于零样本基线
  • 提供高质量人工对齐数据集,适合研究低资源语言

尽管神经机器翻译(NMT)取得进展,提格雷尼亚语等低资源语言仍因语料稀缺、分词策略不足及缺乏标准化评估基准而发展滞后。本文探究使用多语言预训练模型的迁移学习技术,以提升形态丰富、低资源语言的翻译质量。提出一种改进方法,融合语言特异性分词、有据可依的嵌入初始化与领域自适应微调。为实现严格评估,构建了一个高质量、人工对齐的英语-提格雷尼亚语评测数据集,覆盖多个领域。实验结果表明,采用自定义分词的迁移学习显著优于零样本基线,性能提升通过BLEU、chrF及人工定性评估验证。使用邦弗朗尼校正确保配置间统计显著性。错误分析揭示关键局限并指导针对性优化。研究强调语言感知建模与可复现基准在缩小代表性不足语言性能差距中的重要性。资源已公开于 https://github.com/hailaykidu/MachineT_TigEng 及 https://huggingface.co/Hailay/MachineT_TigEng。

原文摘要 · Abstract (English)

Despite advances in Neural Machine Translation (NMT), low-resource languages like Tigrinya remain underserved due to persistent challenges, including limited corpora, inadequate tokenization strategies, and the lack of standardized evaluation benchmarks. This paper investigates transfer learning techniques using multilingual pretrained models to enhance translation quality for morphologically rich, low-resource languages. We propose a refined approach that integrates language-specific tokenization, informed embedding initialization, and domain-adaptive fine-tuning. To enable rigorous assessment, we construct a high-quality, human-aligned English-Tigrinya evaluation dataset covering diverse domains. Experimental results demonstrate that transfer learning with a custom tokenizer substantially outperforms zero-shot baselines, with gains validated by BLEU, chrF, and qualitative human evaluation. Bonferroni correction is applied to ensure statistical significance across configurations. Error analysis reveals key limitations and informs targeted refinements. This study underscores the importance of linguistically aware modeling and reproducible benchmarks in bridging the performance gap for underrepresented languages. Resources are available at https://github.com/hailaykidu/MachineT_TigEng and https://huggingface.co/Hailay/MachineT_TigEng

机器翻译低资源语言多语言模型自定义分词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。