arXiv:2505.11683cs.CL2025-05ACL被引 7

提出改进的双编码器实体消歧模型,性能达新标杆。

Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation

  • 采用上下文标签描述与硬负样本策略提升表示能力
  • 在AIDA-Yago上实现新SOTA,在ZELDA上超越现有方法
  • 适合关注知识库链接与模型设计的NLP研究者

实体消歧(ED)是将文本中的提及链接到知识库条目的任务。双编码器通过将提及和候选标签嵌入共享空间并使用相似度度量进行预测来解决该问题。本文评估了双编码器方法中的关键设计选择,包括损失函数、相似度度量、标签表述格式和负采样策略。提出模型VerbalizED,为文档级双编码器模型,引入上下文相关的标签表述和高效的硬负样本采样机制。此外,还探索了一种迭代预测变体,旨在提升困难样本的消歧效果。在AIDA-Yago上的全面实验验证了方法的有效性,为关键设计选择提供了洞见,并在ZELDA基准上实现了新的状态领先性能。

原文摘要 · Abstract (English)

Entity disambiguation (ED) is the task of linking mentions in text to corresponding entries in a knowledge base. Dual Encoders address this by embedding mentions and label candidates in a shared embedding space and applying a similarity metric to predict the correct label. In this work, we focus on evaluating key design decisions for Dual Encoder-based ED, such as its loss function, similarity metric, label verbalization format, and negative sampling strategy. We present the resulting model VerbalizED, a document-level Dual Encoder model that includes contextual label verbalizations and efficient hard negative sampling. Additionally, we explore an iterative prediction variant that aims to improve the disambiguation of challenging data points. Comprehensive experiments on AIDA-Yago validate the effectiveness of our approach, offering insights into impactful design choices that result in a new State-of-the-Art system on the ZELDA benchmark.

实体消歧双编码器知识库链接

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。