构建首个英-梵语跨语言检索基准,助力古籍数字化与哲学传播
Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents
- 用三种方法融合嵌入与翻译,提升英→梵检索性能
- 翻译法(DT)在3400对数据上最优,超越直接与查询翻译法
- 适合古籍数字化、跨语言AI研究者使用
本研究针对《薄伽梵往世书》章节,构建了一个用于英文查询检索梵文文档的综合性基准。采用三类方法:直接检索(DR)、基于翻译的检索(DT)和查询翻译(QT),结合共享嵌入空间与先进翻译技术,在RAG框架中优化跨语言检索。通过微调前沿模型以适应梵文语言特点,评估了BM25、REPLUG、mDPR、ColBERT、Contriever和GPT-2等模型,并改进梵文摘要技术以支持问答处理。实验表明,DT方法在应对古语文本跨语言挑战时表现最佳。研究基于3,400个英-梵查询-文档对,旨在保护梵文经典并广泛传播其哲学价值。数据集已公开于https://huggingface.co/datasets/manojbalaji1/anveshana。
原文摘要 · Abstract (English)
The study presents a comprehensive benchmark for retrieving Sanskrit documents using English queries, focusing on the chapters of the Srimadbhagavatam. It employs a tripartite approach: Direct Retrieval (DR), Translation-based Retrieval (DT), and Query Translation (QT), utilizing shared embedding spaces and advanced translation methods to enhance retrieval systems in a RAG framework. The study fine-tunes state-of-the-art models for Sanskrit's linguistic nuances, evaluating models such as BM25, REPLUG, mDPR, ColBERT, Contriever, and GPT-2. It adapts summarization techniques for Sanskrit documents to improve QA processing. Evaluation shows DT methods outperform DR and QT in handling the cross-lingual challenges of ancient texts, improving accessibility and understanding. A dataset of 3,400 English-Sanskrit query-document pairs underpins the study, aiming to preserve Sanskrit scriptures and share their philosophical importance widely. Our dataset is publicly available at https://huggingface.co/datasets/manojbalaji1/anveshana
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。