arXiv:2606.06038cs.CL2026-06

用多语言迁移让英语翻译成古普拉克里特语,效果优于基线。

English-to-Prakrit Machine Translation via Multilingual Transfer Learning

  • 通过映射到印地语标签实现古普拉克里特语的低资源翻译
  • 在20样本测试集上达到可比基线的BLEU提升
  • 适合对古典语言翻译和多语言模型迁移感兴趣的研究者

我们在低资源环境下研究英语到普拉克里特语的机器翻译,目标语言未被IndicTrans2支持。通过将普拉克里特语映射至印地语标签(hin_Deva),不修改分词器、词汇表或架构,利用包含1,474组句子的马哈拉施特拉普拉克里特语平行语料库,并在20个样本的阿德拉玛格迪语测试集上进行评估。实验表明,该方法在不调整模型结构的前提下实现了相比未调优基线的语料库BLEU提升。结果说明,基于脚本兼容的语言路由可在无支持语言中实现可行迁移,但受限于数据稀缺与方言差异。代码与训练模型已公开:https://github.com/D3v1s0m/indictrans2-prakrit-mt。

原文摘要 · Abstract (English)

We study English-to-Prakrit machine translation in a low-resource setting where the target language is unsupported by IndicTrans2. We adapt the multilingual model by mapping Prakrit to the Hindi language tag (hin_Deva) without modifying the tokenizer, vocabulary, or architecture. Using a 1,474-pair Maharashtri Prakrit parallel corpus and evaluation on a 20-sample Ardhamagadhi test set, we report corpus BLEU improvements over an untuned baseline. The results indicate that script-compatible language routing can enable feasible transfer to unsupported classical languages, while highlighting limitations due to data scarcity and dialect mismatch. Our code and trained models are released to the public for further exploration https://github.com/D3v1s0m/indictrans2-prakrit-mt.

机器翻译低资源古典语言多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。