用时空一致性提升古意大利语实体链接准确率
DELICATE: Diachronic Entity LInking using Classes And Temporal Evidence
- 结合BERT与Wikidata上下文,通过时间合理性和类型一致判断最佳实体
- 在历史意大利语数据集上超越大模型,尤其对罕见实体表现更优
- 结果可解释性强,适合需要透明推理的数字人文研究
尽管自然语言处理取得显著进展,人文领域中的实体链接(EL)仍面临挑战,主要源于文本类型复杂、缺乏领域专用数据集与模型,以及长尾实体(即知识库中代表性不足的实体)。本文提出两个主要贡献:一是DELICATE,一种新型神经符号方法,用于历史意大利语的实体链接,结合BERT编码器与Wikidata的上下文信息,利用时间合理性与实体类型一致性筛选知识库实体;二是ENEIDE,一个从19至20世纪两部标注文献中半自动构建的多领域历史意大利语文本语料库,涵盖文学与政治文本。实验表明,即便相比参数量达数十亿的大规模模型,DELICATE在历史意大利语上的表现仍更优。进一步分析显示,其置信度分数与特征敏感性使结果比纯神经方法更具可解释性与可读性。
原文摘要 · Abstract (English)
In spite of the remarkable advancements in the field of Natural Language Processing, the task of Entity Linking (EL) remains challenging in the field of humanities due to complex document typologies, lack of domain-specific datasets and models, and long-tail entities, i.e., entities under-represented in Knowledge Bases (KBs). The goal of this paper is to address these issues with two main contributions. The first contribution is DELICATE, a novel neuro-symbolic method for EL on historical Italian which combines a BERT-based encoder with contextual information from Wikidata to select appropriate KB entities using temporal plausibility and entity type consistency. The second contribution is ENEIDE, a multi-domain EL corpus in historical Italian semi-automatically extracted from two annotated editions spanning from the 19th to the 20th century and including literary and political texts. Results show how DELICATE outperforms other EL models in historical Italian even if compared with larger architectures with billions of parameters. Moreover, further analyses reveal how DELICATE confidence scores and features sensitivity provide results which are more explainable and interpretable than purely neural methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。