用AI推荐古典梵文箴言,让古老智慧更易读懂
Pragya: An AI-Based Semantic Recommendation System for Sanskrit Subhasitas
- 用IndicBERT检索相关诗句,再用Mistral生成翻译解释
- 语义检索准确率远超关键词匹配,用户测试满意度高
- 适合研究梵文文化或对AI赋能传统文化感兴趣者
梵文箴言承载千年文化和哲学智慧,但在数字时代因语言和语境障碍而未被充分利用。本文提出Pragya,一种基于检索增强生成(RAG)的梵文箴言语义推荐框架。我们构建了一个包含200句诗的标注数据集,主题标签包括激励、友谊、慈悲等。利用IndicBERT生成句子嵌入,系统可检索与用户查询最相关的前k条诗句;随后将结果输入Mistral LLM模型,生成音译、翻译及上下文解释。实验表明,语义检索在精度和相关性上显著优于关键词匹配;用户研究也证实生成摘要有效提升了可读性。据我们所知,这是首次将检索与生成结合用于梵文箴言的尝试,实现了文化遗产与现代AI应用的融合。
原文摘要 · Abstract (English)
Sanskrit Subhasitas encapsulate centuries of cultural and philosophical wisdom, yet remain underutilized in the digital age due to linguistic and contextual barriers. In this work, we present Pragya, a retrieval-augmented generation (RAG) framework for semantic recommendation of Subhasitas. We curate a dataset of 200 verses annotated with thematic tags such as motivation, friendship, and compassion. Using sentence embeddings (IndicBERT), the system retrieves top-k verses relevant to user queries. The retrieved results are then passed to a generative model (Mistral LLM) to produce transliterations, translations, and contextual explanations. Experimental evaluation demonstrates that semantic retrieval significantly outperforms keyword matching in precision and relevance, while user studies highlight improved accessibility through generated summaries. To our knowledge, this is the first attempt at integrating retrieval and generation for Sanskrit Subhasitas, bridging cultural heritage with modern applied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。