arXiv:2604.26456cs.CLcs.AI2026-04被引 1

构建首个大规模合成梵语命名实体数据集,提升古典文本标注质量。

Naamah: A Large Scale Synthetic Sanskrit NER Corpus via DBpedia Seeding and LLM Generation

  • 结合DBpedia实体提取与240亿参数混合推理模型生成语法自然的文本
  • 产出102,942句高质量银标准梵语语料,支持下游任务训练
  • 适用于梵语自然语言处理研究者及古典文本数字化项目

古典梵语文本的数字化因标注资源稀缺而受阻,尤其是命名实体识别领域。尽管近期方法利用通用大语言模型进行数据增强,但这些方法仍易出错,且缺乏古典语法所需的推理深度。本文提出Naamah,一个包含102,942句子的高质量银标准梵语命名实体数据集。该方法结合DBpedia的实体提取与240亿参数混合推理模型的生成能力,生成语法自然、语义多样化的合成训练数据。我们使用该数据集对两种Transformer架构进行基准测试:多语言巨量模型XLM RoBERTa与参数高效模型IndicBERTv2。

原文摘要 · Abstract (English)

The digitisation of classical Sanskrit literature is impeded by a scarcity of annotated resources, particularly for Named Entity Recognition. While recent methodologies utilise generic Large Language Models (LLMs) for data augmentation, these approaches remain prone to error and often lack the reasoning depth required for classical grammar. In this work, we introduce Naamah, a high quality silver standard Sanskrit NER dataset comprising 102,942 sentences. We propose a methodology that combines entity extraction from DBpedia with the generative capabilities of a 24B parameter hybrid reasoning model to create grammatically natural and synthetically diverse training data. We utilize this dataset to benchmark two transformer architectures: the massive multilingual XLM RoBERTa and the parameter efficient IndicBERTv2.

梵语处理命名实体识别合成数据LLM生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。