构建首个聚焦技术教育领域的印地语多语言翻译数据集与模型
Shiksha: A Technical Domain focused Translation Dataset and Model for Indian Languages
- 通过挖掘NPTEL课程的真人翻译字幕,构建280万条高质量双语对
- 在专业领域任务中超越所有公开模型,通用任务上平均提升超2 BLEU
- 适合低资源印度语言翻译研究者及教育科技开发者使用
神经机器翻译模型通常在科学、技术和教育领域数据上训练不足,导致在涉及科学理解或专业术语的任务中表现不佳,尤其在低资源印度语言上更差。本文通过挖掘NPTEL视频讲座的人工翻译字幕,构建了一个包含超过280万行英文-印地语及印地语-印地语平行语料的多语言语料库,覆盖8种印度语言。我们在此语料库上微调并评估了NMT模型,在领域内任务中优于所有现有公开模型;同时在跨领域任务中,于Flores+基准上平均提升超过2 BLEU。相关模型与数据集已通过Hugging Face发布。
原文摘要 · Abstract (English)
Neural Machine Translation (NMT) models are typically trained on datasets with limited exposure to Scientific, Technical and Educational domains. Translation models thus, in general, struggle with tasks that involve scientific understanding or technical jargon. Their performance is found to be even worse for low-resource Indian languages. Finding a translation dataset that tends to these domains in particular, poses a difficult challenge. In this paper, we address this by creating a multilingual parallel corpus containing more than 2.8 million rows of English-to-Indic and Indic-to-Indic high-quality translation pairs across 8 Indian languages. We achieve this by bitext mining human-translated transcriptions of NPTEL video lectures. We also finetune and evaluate NMT models using this corpus and surpass all other publicly available models at in-domain tasks. We also demonstrate the potential for generalizing to out-of-domain translation tasks by improving the baseline by over 2 BLEU on average for these Indian languages on the Flores+ benchmark. We are pleased to release our model and dataset via this link: https://huggingface.co/SPRINGLab.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。