arXiv:2608.29324cs.CL2026-08

首个基于印度哲学体系的梵语细粒度命名实体识别基准

Padārtha: Ontology-Grounded Fine-Grained NER Benchmark for Classical Sanskrit

论文配图:Padārtha: Ontology-Grounded Fine-Grained NER Benchmark for Classical Sanskrit
图 1 · 摘自论文原文
  • 用古典印度本体论构建标签体系,确保文化语境适配
  • 标注10万+实体提及,含5000句专家验证测试集
  • 生成式模型表现媲美专用模型,但对罕见实体仍不足

标注体系并非中立。将现代新闻文本的标签体系用于古典文学时,会强加源文化定义。本文提出基于《尼耶亚-瓦伊舍提卡》哲学体系的梵语本体论基础细粒度命名实体识别基准「Pad=artha」,以《摩诃婆罗多》为语料。标签体系包含18个细粒度类别,归入10个本体节点,映射至5个标准粗粒度标签,保证与现有基准互操作性。专家标注了超过12.6千条来自学术索引的实体条目,对应《玛哈纳马》语料库中的108,335个实体提及,覆盖73,632行诗句,并构建了5,000行诗句的专家验证测试集,重点覆盖稀有提及。首次系统对比生成式与传统架构在梵语上的表现,发现微调后的生成式模型性能可媲美专用系统。但所有系统在从粗到细粒度时均出现显著下降,且对训练中未见的实体提及表现差。问题不单是数据稀缺,因为微调模型对未见实体的召回远低于已见实体,且在词汇歧义下倾向于默认多数语义。

原文摘要 · Abstract (English)

Annotation schemas are not neutral. When applied to classical literature, tag sets developed for modern journalistic texts impose source-culture definitions on texts they were never designed to describe. We instead ground a schema in the tradition of the text itself introducing \textit{Padārtha}, the first ontology-grounded fine-grained Named Entity Recognition (NER) benchmark for Sanskrit, built on the \textit{Mahābhārata} epic. Our tag set derives from \textit{Nyāya-Vaiśesika}, a classical Indian ontological system, yielding 18 fine-grained categories organized under 10 ontological nodes and mapped onto five standard coarse tags, ensuring interoperability with existing benchmarks. Expert annotators label over 12.6K entries from a scholarly index of named entities, linked to corresponding mentions in the \textit{Mahānāma} corpus, producing fine-grained annotations for 108,335 entity mentions across 73,632 verses, along with a 5,000-verse expert-verified test set sampled to stress rare mentions. We present the first systematic benchmarking of generative NER against traditional architectures for Sanskrit, finding that fine-tuned generative models perform comparably to task-specific systems. However, all systems show a sharp decline from coarse to fine granularity and struggle with out-of-entity mentions unseen during training. The limitation is not due to data scarcity alone, as fine-tuned models recall unseen entities far worse than seen ones and tend to default to the majority sense under lexical ambiguity.

命名实体识别梵语处理本体论细粒度标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。