统一多语言知识图谱的文本与关系补全,提升信息完整性。
KG-TRICK: Unifying Textual and Relational Information Completion of Knowledge for Multilingual Knowledge Graphs
- 设计序列到序列框架,联合补全知识图谱的文本与关系信息。
- 在10种语言超过2.5万实体上验证,显著提升补全效果。
- 适合多语言知识库构建与增强的研究者使用。
多语言知识图谱(KG)为多种自然语言处理应用提供高质量的关系与文本信息,但普遍存在不完整问题,尤其在非英语语言中更为严重。以往研究将知识图谱补全(KGC)和知识图谱增强(KGE)视为独立任务,分别预测缺失关系或实体文本信息。我们提出二者存在内在关联且可互为助力。为此,本文引入KG-TRICK,一种统一文本与关系信息补全的序列到序列新框架。实验表明:1)可将KGC与KGE统一于同一框架;2)融合多语言文本信息能有效提升知识图谱的完整性。作为贡献之一,我们还构建了目前最大规模的人工标注基准WikiKGE10++,涵盖10种语言、超2.5万实体。
原文摘要 · Abstract (English)
Multilingual knowledge graphs (KGs) provide high-quality relational and textual information for various NLP applications, but they are often incomplete, especially in non-English languages. Previous research has shown that combining information from KGs in different languages aids either Knowledge Graph Completion (KGC), the task of predicting missing relations between entities, or Knowledge Graph Enhancement (KGE), the task of predicting missing textual information for entities. Although previous efforts have considered KGC and KGE as independent tasks, we hypothesize that they are interdependent and mutually beneficial. To this end, we introduce KG-TRICK, a novel sequence-to-sequence framework that unifies the tasks of textual and relational information completion for multilingual KGs. KG-TRICK demonstrates that: i) it is possible to unify the tasks of KGC and KGE into a single framework, and ii) combining textual information from multiple languages is beneficial to improve the completeness of a KG. As part of our contributions, we also introduce WikiKGE10++, the largest manually-curated benchmark for textual information completion of KGs, which features over 25,000 entities across 10 diverse languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。