零样本对话实体链接在新领域知识库上表现差,需新评估方法。
Real World Conversational Entity Linking Requires More Than Zeroshots
- 构建基于Reddit的零样本对话实体链接数据集,测试模型跨知识库泛化能力。
- 在新领域知识库Fandom上,零样本模型性能显著下降,准确率大幅降低。
- 研究揭示现有评估方法不足,适合资源受限场景下的对话系统开发者参考。
对话中的实体链接(EL)在实际应用中面临显著挑战,主要源于实体标注的对话数据集稀缺以及包含特定领域、长尾实体的知识库稀疏。我们设计了针对性的评估场景,以衡量资源受限下EL模型的有效性。评估使用两个知识库:体现真实世界复杂性的Fandom和广泛使用的Wikipedia。首先,我们通过基于Reddit讨论构建的新型零样本对话实体链接数据集,评估模型在未见过的新知识库(Fandom)上的泛化能力;其次,评估模型在无预先训练的情况下适应对话场景的能力。结果表明,当前零样本EL模型在未预先训练且引入新领域知识库时性能显著下降。研究发现,以往的评估方法未能充分捕捉零样本EL的真实复杂性,凸显了设计与评估对话式EL模型以应对资源有限情况的必要性。本研究提出的评估设置和数据集已公开可用。
原文摘要 · Abstract (English)
Entity linking (EL) in conversations faces notable challenges in practical applications, primarily due to the scarcity of entity-annotated conversational datasets and sparse knowledge bases (KB) containing domain-specific, long-tail entities. We designed targeted evaluation scenarios to measure the efficacy of EL models under resource constraints. Our evaluation employs two KBs: Fandom, exemplifying real-world EL complexities, and the widely used Wikipedia. First, we assess EL models' ability to generalize to a new unfamiliar KB using Fandom and a novel zero-shot conversational entity linking dataset that we curated based on Reddit discussions on Fandom entities. We then evaluate the adaptability of EL models to conversational settings without prior training. Our results indicate that current zero-shot EL models falter when introduced to new, domain-specific KBs without prior training, significantly dropping in performance. Our findings reveal that previous evaluation approaches fall short of capturing real-world complexities for zero-shot EL, highlighting the necessity for new approaches to design and assess conversational EL models to adapt to limited resources. The evaluation setup and the dataset proposed in this research are made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。