arXiv:2503.21530cs.CLcs.AI2025-03被引 5

用Transformer模型提升罗马乌尔都语与乌尔都语间的低资源翻译质量。

Low-Resource Transliteration for Roman-Urdu and Urdu Using Transformer-Based Models

  • 基于m2m100模型,结合掩码语言建模预训练与双数据集微调。
  • 罗马语转乌尔都语、乌尔都语转罗马语的字符级BLEU分别达96.37和97.44。
  • 克服旧方法评估缺陷,适合低资源语言文本转换任务研究者。

随着信息检索领域日益重视包容性,低资源语言的需求仍面临重大挑战。尽管南亚地区广泛使用乌尔都语及其罗马化形式罗马乌尔都语,但两者之间的音译研究仍不充分。以往基于RNN在Roman-Urdu-Parl数据集上的工作虽表现良好,但存在领域适应性差和评估不严谨的问题。本文提出一种基于m2m100多语言翻译模型的Transformer方法,通过掩码语言建模(MLM)预训练,并在Roman-Urdu-Parl和领域多样化的Dakshina数据集上进行微调。为解决以往评估缺陷,引入严格的分数据集划分,并采用BLEU、字符级BLEU和CHRF进行性能评估。模型取得优异结果:乌尔都语→罗马乌尔都语的字符级BLEU达96.37,罗马乌尔都语→乌尔都语达97.44,优于此前的RNN基线及GPT-4o Mini,验证了多语言迁移学习在低资源音译任务中的有效性。

原文摘要 · Abstract (English)

As the Information Retrieval (IR) field increasingly recognizes the importance of inclusivity, addressing the needs of low-resource languages remains a significant challenge. Transliteration between Urdu and its Romanized form, Roman Urdu, remains underexplored despite the widespread use of both scripts in South Asia. Prior work using RNNs on the Roman-Urdu-Parl dataset showed promising results but suffered from poor domain adaptability and limited evaluation. We propose a transformer-based approach using the m2m100 multilingual translation model, enhanced with masked language modeling (MLM) pretraining and fine-tuning on both Roman-Urdu-Parl and the domain-diverse Dakshina dataset. To address previous evaluation flaws, we introduce rigorous dataset splits and assess performance using BLEU, character-level BLEU, and CHRF. Our model achieves strong transliteration performance, with Char-BLEU scores of 96.37 for Urdu->Roman-Urdu and 97.44 for Roman-Urdu->Urdu. These results outperform both RNN baselines and GPT-4o Mini and demonstrate the effectiveness of multilingual transfer learning for low-resource transliteration tasks.

音译低资源Transformer多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。