arXiv:2601.06400cs.CL2026-01被引 7

构建古佛教文献多语言平行语料库与专用大模型,助力跨语言文本分析。

MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan

  • 设计多语言并行句对挖掘流程,自动识别梵文、汉文、藏文等古语文本对应关系
  • 建成包含174万句对的大型平行语料库,支持四类经典语言间的机器翻译与语义检索
  • 推出Gemma 2 MITRA系列模型,在翻译和语义嵌入任务上均达顶尖水平,适合佛学与古典文献研究

古代佛教文献中频繁出现跨语言的文本对应,涉及梵文、巴利文、佛教汉语、藏文等多种语言,但往往缺乏标注。这类材料规模庞大,人工分析难以实现。我们提出MITRA框架,包括一种新型多语言并行段落挖掘方法MITRA-parallel,构建了涵盖174万句对的大型平行语料库(覆盖梵文、中文、藏文),并开发了领域专用预训练语言模型Gemma 2 MITRA。其中,Gemma 2 MITRA-MT在机器翻译任务上达到当前最优性能,优于许多更大的开源模型;Gemma 2 MITRA-E作为语义嵌入模型,在新构建的细粒度语义相似性基准上表现领先。相关平行数据集、模型权重及评测基准均已开源,可支持自然语言处理与佛教及古典亚洲文学的学术研究。

原文摘要 · Abstract (English)

Ancient Buddhist literature features frequent, yet often unannotated, textual parallels spread across diverse languages: Sanskrit, Pāli, Buddhist Chinese, Tibetan, and more. The scale of this material makes manual examination prohibitive. We present the MITRA framework, which consists of a novel pipeline for multilingual parallel passage mining, MITRA-parallel, a large-scale corpus of 1.74 million parallel sentence pairs between Sanskrit, Chinese, and Tibetan, and the development of the domain-specific pretrained language model Gemma 2 MITRA. We present Gemma 2 MITRA-MT, a version of this base model fine-tuned on machine translation tasks, reaching state-of-the-art performance for machine translation of these languages into English and outperforming even much larger open-source models. We also present Gemma 2 MITRA-E, a semantic embedding model that shows state-of-the-art performance on a novel, detailed semantic embedding benchmark. We make the parallel dataset, model weights, and semantic similarity benchmark openly available to aid both NLP research and philological studies in Buddhist and classical Asian literature.

古语文本多语言语义检索佛教文献

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。