arXiv:2607.05849cs.CL2026-07

利用西里尔字母作桥梁,提升传统蒙古文翻译准确率

CoPiT: Cognitive Pivot Translation for Digraphic Low-Resource Mongolian in the Traditional Script

论文配图:CoPiT: Cognitive Pivot Translation for Digraphic Low-Resource Mongolian in the Traditional Script
图 1 · 摘自论文原文
  • 通过西里尔字母作为中间桥接,解决传统蒙古文拼写歧义问题
  • 相比直接翻译,BLEU提升显著,COMET得分提高1.5至1.6倍
  • 适合低资源语言翻译研究者,尤其关注蒙古文与多文字系统

低资源语言的机器翻译仍具挑战性,蒙古语是典型代表。作为双文字语言,蒙古语使用西里尔字母和传统文字两种书写系统,数据分布严重失衡:西里尔字母语料丰富,而传统文字语料极度稀缺且拼写模糊,导致直接翻译性能大幅下降。本文提出CoPiT,一种受认知启发的基于中间桥接的翻译流程,利用内部资源层级,将翻译路径经由西里尔字母进行。该流程在翻译前显式化解传统文字中的脚本歧义,实现更稳定准确的意义传递。在多个主干模型和目标语言上,CoPiT均优于直接翻译,取得显著的绝对BLEU提升,并带来1.5-1.6倍的COMET增益。这些改进使开源模型在相似评估条件下达到或超越GPT-4.1表现。此外,CoPiT可直接从传统文字文本生成合成平行语料,缓解真实低资源场景下的数据匮乏问题。我们发布了涵盖蒙古语(两种书写系统)及英语、韩语、俄语的新多文字平行数据集。所有数据集与代码已公开于 https://anonymous.4open.science/r/anonymous_project-76C7。

原文摘要 · Abstract (English)

Low-resource languages remain challenging for machine translation, and Mongolian is a representative case. As a digraphic language, Mongolian is written in both Cyrillic and Traditional scripts, which exhibit a severe imbalance in data availability. While the Cyrillic script is relatively well-resourced, the Traditional script remains extremely data-scarce and orthographically ambiguous, leading to substantial performance degradation in direct translation. We propose CoPiT, a cognitively motivated pivot-based translation pipeline that exploits this internal resource hierarchy by routing translation through the Cyrillic script. The pipeline explicitly resolves script-induced ambiguity in the Traditional script before translation, enabling more stable and accurate meaning transfer. Across multiple backbone models and target languages, CoPiT consistently outperforms direct translation, achieving substantial absolute BLEU improvements together with consistent 1.5-1.6x COMET gains. These gains allow strong open-source models to match or outperform GPT-4.1 under comparable evaluation settings. Beyond inference-time improvements, CoPiT enables the construction of synthetic parallel data directly from Traditional-script text, mitigating data scarcity in realistic low-resource scenarios. We release a new multi-script parallel dataset covering Mongolian in both scripts alongside English, Korean, and Russian. All datasets and code are publicly available at https://anonymous.4open.science/r/anonymous_project-76C7.

机器翻译低资源语言双文字系统蒙古语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。