arXiv:2409.07869cs.CL2024-09

用大模型辅助规则学习,提升知识图谱补全质量

Learning Rules from KGs Guided by Language Models

  • 结合大模型预测结果优化规则排序
  • 在不完整/有偏图谱中显著提升规则准确性
  • 适合做知识图谱补全的研究者参考

信息抽取技术推动了大规模知识图谱(如Yago、Wikidata或Google KG)的自动构建,广泛应用于语义搜索和数据分析。然而,由于半自动化构建,知识图谱常存在不完整性。规则学习方法通过从图谱中提取频繁模式并转化为规则,可用于预测缺失事实。其关键步骤是规则排序。在高度不完整或有偏的知识图谱(如主要包含名人事实)中,传统统计指标(如规则置信度)易使有偏规则被错误地排在前列。此前研究提出结合知识图谱嵌入模型预测的事实进行规则排序。随着大语言模型(LMs)兴起,有研究声称其可替代性用于知识图谱补全。本文旨在验证大模型在提升规则学习系统质量方面的实际作用。

原文摘要 · Abstract (English)

Advances in information extraction have enabled the automatic construction of large knowledge graphs (e.g., Yago, Wikidata or Google KG), which are widely used in many applications like semantic search or data analytics. However, due to their semi-automatic construction, KGs are often incomplete. Rule learning methods, concerned with the extraction of frequent patterns from KGs and casting them into rules, can be applied to predict potentially missing facts. A crucial step in this process is rule ranking. Ranking of rules is especially challenging over highly incomplete or biased KGs (e.g., KGs predominantly storing facts about famous people), as in this case biased rules might fit the data best and be ranked at the top based on standard statistical metrics like rule confidence. To address this issue, prior works proposed to rank rules not only relying on the original KG but also facts predicted by a KG embedding model. At the same time, with the recent rise of Language Models (LMs), several works have claimed that LMs can be used as alternative means for KG completion. In this work, our goal is to verify to which extent the exploitation of LMs is helpful for improving the quality of rule learning systems.

知识图谱规则学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。