arXiv:2511.00486cs.CL2025-11EMNLP

构建首个跨语言的比利语-印地语-英语平行语料库,推动小语种机器翻译研究

Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus

  • 创建包含11万句的多语言语料库,覆盖教育、政务等关键领域
  • 微调NLLB-200轻量模型在比利语翻译中表现最佳,准确率显著提升
  • 验证大模型在少资源语言上的泛化能力,适合关注包容性AI的研究者

印度语言多样性带来巨大机器翻译挑战,尤其对缺乏高质量语言资源的部落语言如比利语。本文提出首个全球最大的比利语-印地语-英语平行语料库(BHEPC),包含11万条经专家校对的句子,涵盖教育、行政与新闻等关键领域,为低资源机器翻译研究提供重要基准。为建立全面的比利语机器翻译评估体系,我们在英-印地-比利三语双向翻译任务上测试了多种开源与专有多语言大模型。实验表明,微调后的NLLB-200轻量化变体(600M参数)表现最优。此外,通过上下文学习方法评估多语言大模型在BHEPC上的生成翻译能力,分析其跨领域泛化性能并量化分布偏差。该工作填补关键资源空白,推动面向低资源与边缘语言的包容性自然语言处理技术发展。

原文摘要 · Abstract (English)

The linguistic diversity of India poses significant machine translation challenges, especially for underrepresented tribal languages like Bhili, which lack high-quality linguistic resources. This paper addresses the gap by introducing Bhili-Hindi-English Parallel Corpus (BHEPC), the first and largest parallel corpus worldwide comprising 110,000 meticulously curated sentences across Bhili, Hindi, and English. The corpus was created with the assistance of expert human translators. BHEPC spans critical domains such as education, administration, and news, establishing a valuable benchmark for research in low resource machine translation. To establish a comprehensive Bhili Machine Translation benchmark, we evaluated a wide range of proprietary and open-source Multilingual Large Language Models (MLLMs) on bidirectional translation tasks between English/Hindi and Bhili. Comprehensive evaluation demonstrates that the fine-tuned NLLB-200 distilled 600M variant model outperforms others, highlighting the potential of multilingual models in low resource scenarios. Furthermore, we investigated the generative translation capabilities of multilingual LLMs on BHEPC using in-context learning, assessing performance under cross-domain generalization and quantifying distributional divergence. This work bridges a critical resource gap and promotes inclusive natural language processing technologies for low-resource and marginalized languages globally.

低资源翻译多语言模型语料库构建小语种

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。