构建首个面向日本实体链接的标注语料库,填补日文资源空白。
CADEL: A Corpus of Administrative Web Documents for Japanese Entity Linking
- 基于规范设计策略,构建覆盖日本特有实体的日语文本语料
- 标注一致性高,含大量需消歧的复杂案例
- 适合日文信息抽取与知识图谱研究者使用
实体链接是将语言表达式与代表现实世界实体和概念的知识库条目关联的任务。现有语言资源主要针对英语,可用于评估日文系统资源仍有限。本研究提出实体链接语料库的设计策略,并构建一个用于训练和评估日文实体链接系统的标注语料库,涵盖大量与日本相关的特定实体表达。评估显示标注者间一致性高,基于字符串匹配的初步消歧实验表明语料包含大量非平凡案例,支持其作为评估基准的潜力。
原文摘要 · Abstract (English)
Entity linking is the task of associating linguistic expressions with entries in a knowledge base that represent real-world entities and concepts. Language resources for this task have primarily been developed for English, and the resources available for evaluating Japanese systems remain limited. In this study, we develop a corpus design policy for the entity linking task and construct an annotated corpus for training and evaluating Japanese entity linking systems, with rich coverage of linguistic expressions referring to entities that are specific to Japan. Evaluation of inter-annotator agreement confirms the high consistency of the annotations in the corpus, and a preliminary experiment on entity disambiguation based on string matching suggests that the corpus contains a substantial number of non-trivial cases, supporting its potential usefulness as an evaluation benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。