arXiv:2509.19844cs.CL2025-09中稿 · EMNLP

首个梵语文学实体识别与链接数据集,助力复杂文本实体消解研究

Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking

  • 构建首个梵语史诗级文学文本的端到端实体发现与链接数据集
  • 包含10.9万实体提及对应5500个唯一实体,支持跨语言知识库对齐
  • 适用于文学文本、低资源语言及长距离指代消解等研究方向

高词汇变体、指代模糊和长距离依赖使文学文本中的实体消解尤为困难。我们提出Mahānāma,首个针对梵语这一形态丰富且资源匮乏语言的大规模端到端实体发现与链接(EDL)数据集。该数据集源自《摩诃婆罗多》——世界最长史诗,包含超过10.9万条命名实体提及,对应5500个唯一实体,并与英文知识库对齐以支持跨语言链接。其复杂的叙事结构、广泛的名字变体和指代模糊性,对消解系统构成严峻挑战。评估显示,当前共指消解与实体链接模型在测试集全局上下文下表现不佳,揭示了现有方法在复杂语篇中解决实体问题的局限性。Mahānāma因此成为推动实体消解技术,尤其是在文学领域的重要基准。

原文摘要 · Abstract (English)

High lexical variation, ambiguous references, and long-range dependencies make entity resolution in literary texts particularly challenging. We present Mahānāma, the first large-scale dataset for end-to-end Entity Discovery and Linking (EDL) in Sanskrit, a morphologically rich and under-resourced language. Derived from the Mahābhārata, the world's longest epic, the dataset comprises over 109K named entity mentions mapped to 5.5K unique entities, and is aligned with an English knowledge base to support cross-lingual linking. The complex narrative structure of Mahānāma, coupled with extensive name variation and ambiguity, poses significant challenges to resolution systems. Our evaluation reveals that current coreference and entity linking models struggle when evaluated on the global context of the test set. These results highlight the limitations of current approaches in resolving entities within such complex discourse. Mahānāma thus provides a unique benchmark for advancing entity resolution, especially in literary domains.

实体链接梵语文学分析知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。