构建11种印度语言平行语料库,助力多语言机器翻译研究
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
- 构建涵盖11种印度语言的77.2万句对平行语料库
- 在政府、医疗、通用领域验证模型表现差异
- 适合关注印度语言翻译与跨语言迁移的研究者
印度拥有超过120种主要语言及约1600种其他语言,其中22种为宪法规定的官方语言。尽管多语言神经机器翻译(NMT)取得进展,但高质量印度语言平行语料仍严重缺乏,尤其在不同领域间。本文提出大规模高质量标注平行语料库CorIL,覆盖英语、泰卢固语、印地语、旁遮普语、奥里亚语、克什米尔语、信德语、多格里语、卡纳达语、乌尔都语和古吉拉特语共11种语言,总计77.2万句对。语料按政府、医疗、通用三大领域系统分类,支持领域感知的机器翻译研究与领域适应。通过微调并评估IndicTrans2、NLLB和BhashaVerse等先进NMT模型,揭示了语言书写系统对性能的影响:多语言模型在波斯-阿拉伯字母(乌尔都语、信德语)上表现更优,而其他模型在印度字母系上占优。本文提供详细的领域与书写系统性能分析,为跨脚本迁移学习提供洞见。公开发布CorIL,旨在显著提升印度语言训练数据的可得性。
原文摘要 · Abstract (English)
India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite recent progress in multilingual neural machine translation (NMT), high-quality parallel corpora for Indian languages remain scarce, especially across varied domains. In this paper, we introduce a large-scale, high-quality annotated parallel corpus covering 11 of these languages : English, Telugu, Hindi, Punjabi, Odia, Kashmiri, Sindhi, Dogri, Kannada, Urdu, and Gujarati comprising a total of 772,000 bi-text sentence pairs. The dataset is carefully curated and systematically categorized into three key domains: Government, Health, and General, to enable domain-aware machine translation research and facilitate effective domain adaptation. To demonstrate the utility of CorIL and establish strong benchmarks for future research, we fine-tune and evaluate several state-of-the-art NMT models, including IndicTrans2, NLLB, and BhashaVerse. Our analysis reveals important performance trends and highlights the corpus's value in probing model capabilities. For instance, the results show distinct performance patterns based on language script, with massively multilingual models showing an advantage on Perso-Arabic scripts (Urdu, Sindhi) while other models excel on Indic scripts. This paper provides a detailed domain-wise performance analysis, offering insights into domain sensitivity and cross-script transfer learning. By publicly releasing CorIL, we aim to significantly improve the availability of high-quality training data for Indian languages and provide a valuable resource for the machine translation research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。