构建首个覆盖三千年梵文多领域翻译数据集,助力复杂文本机器翻译研究
Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
- 构建39万+对梵英平行语料,覆盖宗教、哲学、诗歌等多元文本
- 比现有最大数据集大四倍,含精细时间与领域标注,支持细粒度分析
- 适用于研究梵文复杂表达翻译,适合语言学与跨文化研究者
尽管机器翻译在高资源语言中被视为“已解决”,但面对诗歌、哲思、多重隐喻等复杂内容时仍面临挑战。梵文文献正是此类难题的典型代表:兼具音变、复合词、重形态等语言特征,并涵盖从仪式文本到史诗、哲学论著、科学文献等跨越三千年的广泛领域。当前缺乏覆盖多时期与多领域的公开资源。为此,我们推出Mitrasamgraha,一个高质量梵英机器翻译数据集,包含391,548对平行语料,规模超过此前最大数据集Itih=asa的四倍。数据覆盖三千年历史,涵盖多种文体。与网络爬取数据不同,本数据集提供精确的时间与领域标注,支持对翻译性能随时代与领域变化的细粒度研究。我们还发布验证集(5,587对)和测试集(5,552对),均经后校正。实验表明,商用模型与开源模型在该数据集上微调后性能显著提升,但在处理复杂复合词、哲学概念与多重隐喻方面仍存明显挑战。同时分析了上下文学习对商业模型表现的影响。
原文摘要 · Abstract (English)
While machine translation is regarded as a "solved problem" for many high-resource languages, close analysis quickly reveals that this is not the case for content that shows challenges such as poetic language, philosophical concepts, multi-layered metaphorical expressions, and more. Sanskrit literature is a prime example of this, as it combines a large number of such challenges in addition to inherent linguistic features like sandhi, compounding, and heavy morphology, which further complicate NLP downstream tasks. It spans multiple millennia of text production time as well as a large breadth of different domains, ranging from ritual formulas via epic narratives, philosophical treatises, poetic verses up to scientific material. As of now, there is a strong lack of publicly available resources that cover these different domains and temporal layers of Sanskrit. We therefore introduce Mitrasamgraha, a high-quality Sanskrit-to-English machine translation dataset consisting of 391,548 bitext pairs, more than four times larger than the largest previously available Sanskrit dataset Itih=asa. It covers a time period of more than three millennia and a broad range of historical Sanskrit domains. In contrast to web-crawled datasets, the temporal and domain annotation of this dataset enables fine-grained study of domain and time period effects on MT performance. We also release a validation set consisting of 5,587 and a test set consisting of 5,552 post-corrected bitext pairs. We conduct experiments benchmarking commercial and open models on this dataset and fine-tune NLLB and Gemma models on the dataset, showing significant improvements, while still recognizing significant challenges in the translation of complex compounds, philosophical concepts, and multi-layered metaphors. We also analyze how in-context learning on this dataset impacts the performance of commercial models
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。