arXiv:2608.25038cs.CL2026-08中稿 · EMNLP

构建梵语词汇表任务,让模型从诗文对中自动提炼语义短语并生成翻译锚定的释义。

Padamitra: Grounded Glossary Generation for Classical Sanskrit

  • 基于诗句-译文对,设计可评估的梵语词汇生成框架。
  • 指令微调使模型性能显著优于提示工程,分割策略提升准确性。
  • 适合研究古典语言处理与文本注释的学者使用。

我们提出了一种名为「扎根词汇生成」的结构化任务,要求模型从一首梵文诗句及其对应翻译中恢复出语义有意义的梵语短语,并生成以翻译为依据的释义,将传统的帕塔评论实践形式化为可评估的自然语言处理目标。我们构建了一个包含31,316个诗句-翻译-词汇三元组的基准数据集,涵盖《罗摩衍那》和《薄伽梵往世书》,并引入两个评估指标:用于短语恢复的杰卡德相似度,以及用于语义一致性的意义忠实度。在Gemma-3n-E4B、Gemma-3-12B、Phi-4和Qwen3.5-9B等模型的零样本、少样本及指令微调版本中,指令微调显著优于提示方法,而显式分词进一步带来性能提升。错误分析揭示,沙迪(sandhi)和萨马萨(samasa)复合词的过度分割是主要失败模式,表明形态建模是实现忠实梵语词素分解的关键瓶颈。

原文摘要 · Abstract (English)

We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Across zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segmentation yields gains. Error analysis identifies over-segmentation of sandhi and samasa compounds as the dominant failure mode, pointing to morphological modeling as the key bottleneck for faithful Sanskrit lexical decomposition.

古典语言词汇生成梵语处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。