构建跨域通用的塔吉克-波斯文互译模型,突破领域限制。
ParsTranslit: Truly Versatile Tajik-Farsi Transliteration
- 统一多源数据训练序列到序列模型,提升跨领域适应性
- 在双向翻译中分别达到87.91和92.28的chrF++分数
- 开源数据与代码,适合语言处理与多语种应用研究者
波斯语为双书写体系语言,伊朗与阿富汗使用波斯-阿拉伯文,塔吉克斯坦使用塔吉克-西里尔文。尽管两国方言高度相似,但文字差异导致无法直接一一对应,阻碍了书面交流。此前研究多依赖自建数据集,仅适用于古诗或词汇表等特定领域,缺乏实际可用的泛化能力。本文提出首个基于所有可用数据集训练的塔吉克-波斯文互译序列模型,并构建两个新数据集。实验覆盖多种文本领域,清晰揭示任务难度,建立全面可比基准。模型在波斯文到塔吉克文翻译中获得87.91的chrF++和0.05的归一化字符错误率(CER),塔吉克文到波斯文则达92.28和0.04。模型、数据与代码已公开于https://anonymous.4open.science/r/ParsTranslit-FB30/。
原文摘要 · Abstract (English)
As a digraphic language, the Persian language utilizes two written standards: Perso-Arabic in Afghanistan and Iran, and Tajik-Cyrillic in Tajikistan. Despite the significant similarity between the dialects of each country, script differences prevent simple one-to-one mapping, hindering written communication and interaction between Tajikistan and its Persian-speaking ``siblings''. To overcome this, previously-published efforts have investigated machine transliteration models to convert between the two scripts. Unfortunately, most efforts did not use datasets other than those they created, limiting these models to certain domains of text such as archaic poetry or word lists. A truly usable transliteration system must be capable of handling varied domains, meaning that suck models lack the versatility required for real-world usage. The contrast in domain between data also obscures the task's true difficulty. We present a new state-of-the-art sequence-to-sequence model for Tajik-Farsi transliteration trained across all available datasets, and present two datasets of our own. Our results across domains provide clearer understanding of the task, and set comprehensive comparable leading benchmarks. Overall, our model achieves chrF++ and Normalized CER scores of 87.91 and 0.05 from Farsi to Tajik and 92.28 and 0.04 from Tajik to Farsi. Our model, data, and code are available at https://anonymous.4open.science/r/ParsTranslit-FB30/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。