arXiv:2603.24307cs.CL2026-03被引 1

构建首个大规模现代印地-梵语平行语料库,提升机器翻译性能。

Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation

  • 构建涵盖现代口语、儿童读物等多源的9.2万句平行语料
  • 模型在自域测试集上表现显著提升,跨数据集表现稳定
  • 与现有语料重叠少,填补低资源印地语翻译空白

我们发布 Samasāmayik,一个精心构建的大规模印地语-梵语平行语料库,包含92,196对平行句子。不同于以往以古典文献和诗歌为主的梵语文本,该语料库整合了口语教程、儿童杂志、广播对话及说明材料等当代内容。通过微调ByT5、NLLB和IndicTrans-v2三种互补模型,我们验证了该数据集的实用性。实验表明,基于Samasāmayik训练的模型在自域测试集上取得显著性能提升,同时在其他常用测试集上表现相当,确立了现代印地-梵语翻译的新基准。进一步对比分析显示,该语料与现有资源在语义和词汇层面重叠极低,证实其新颖性与非冗余性,是低资源印地语机器翻译的重要新资源。

原文摘要 · Abstract (English)

We release Samasāmayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focuses on classical era text and poetry, this corpus aggregates data from diverse sources covering contemporary materials, including spoken tutorials, children's magazines, radio conversations, and instruction materials. We benchmark this new dataset by fine-tuning three complementary models - ByT5, NLLB and IndicTrans-v2, to demonstrate its utility. Our experiments demonstrate that models trained on the Samasamayik corpus achieve significant performance gains on in-domain test data, while achieving comparable performance on other widely used test sets, establishing a strong new performance baseline for contemporary Hindi-Sanskrit translation. Furthermore, a comparative analysis against existing corpora reveals minimal semantic and lexical overlap, confirming the novelty and non-redundancy of our dataset as a robust new resource for low-resource Indic language MT.

机器翻译梵语印地语低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。