构建高质量印地语-马拉地语翻译数据集,解决低资源语言训练数据匮乏问题。
BhashaSetu: A Data-Centric Approach to Low-Resource Machine Translation

- 整合新闻、医疗、文学等多源数据,构建278万句对的平行语料库。
- 去重处理使翻译性能提升1.17 BLEU和2.21 chrF++,显著优于其他预处理。
- 适合低资源语言研究者、需要形态学敏感模型的NLP开发者使用。
我们提出BhashaSetu,一个面向低资源神经机器翻译的英语-马拉地语平行语料库。马拉地语使用者超9500万,但高质量平行语料稀缺。该数据集包含来自新闻、政治、医疗、文学和文化等多样来源的278万句对,并提供词干化与词形还原表示,支持形态学分析。我们在多个SOTA模型上进行基准测试,采用BLEU、spBLEU、chrF++和TER指标评估。通过LoRA对NLLB-200-distilled-600M进行参数高效微调。关键发现:语料级去重是提升下游性能的最关键预处理步骤(不执行则性能下降1.17 BLEU和2.21 chrF++),表明对多源语料的严谨清洗是低成本高回报的干预手段。数据集已公开,以推动可复现的、语言学驱动的低资源机器翻译研究。
原文摘要 · Abstract (English)
We present BhashaSetu, a linguistically enriched English--Marathi parallel dataset addressing persistent data limitations in low-resource neural machine translation (NMT). Marathi, spoken by over 95 million people, remains underrepresented in high-quality parallel corpora across diverse domains. Our dataset comprises 2.78 million sentence pairs from heterogeneous sources including news, politics, healthcare, literature, and culture, with stemmed and lemmatized representations to support morphology-aware analysis. We benchmark multiple state-of-the-art translation models using BLEU, spBLEU, chrF++, and TER metrics, and conduct parameter-efficient fine-tuning of NLLB-200-distilled-600M using LoRA. A key finding from our ablation: corpus-level deduplication is the single largest preprocessing contributor to downstream quality (removing it reduces performance by 1.17 BLEU and 2.21 chrF++), demonstrating that disciplined cross-source corpus hygiene is a low-cost, high-impact intervention for low-resource, morphologically rich languages. The dataset is publicly released to promote reproducible and linguistically informed low-resource NMT research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。