arXiv:2409.04056cs.AIcs.CL2024-09被引 6

用大模型自动清理维基数据分类体系,提升准确性与可用性。

Refining Wikidata Taxonomy using Large Language Models

  • 结合大模型与图挖掘技术,通过零样本提示自动修正分类关系。
  • 在实体类型识别任务中表现优于原始维基数据,证明实用性。
  • 适合需要高质量知识图谱的开发者与研究者使用。

由于其协作性质,维基数据的分类体系复杂,存在实例与类别的混淆、部分分类路径不准确、循环结构以及类之间高度冗余等问题。手动清理耗时且易出错或带有主观判断。我们提出WiKC,一种利用大语言模型(LLMs)与图挖掘技术自动清理的维基数据新版本分类体系。通过在开源大模型上使用零样本提示,执行剪枝链接或合并类别等操作。从内在与外在两个角度评估了优化后分类体系的质量,后者基于实体类型识别任务,验证了WiKC的实际价值。

原文摘要 · Abstract (English)

Due to its collaborative nature, Wikidata is known to have a complex taxonomy, with recurrent issues like the ambiguity between instances and classes, the inaccuracy of some taxonomic paths, the presence of cycles, and the high level of redundancy across classes. Manual efforts to clean up this taxonomy are time-consuming and prone to errors or subjective decisions. We present WiKC, a new version of Wikidata taxonomy cleaned automatically using a combination of Large Language Models (LLMs) and graph mining techniques. Operations on the taxonomy, such as cutting links or merging classes, are performed with the help of zero-shot prompting on an open-source LLM. The quality of the refined taxonomy is evaluated from both intrinsic and extrinsic perspectives, on a task of entity typing for the latter, showing the practical interest of WiKC.

知识图谱大模型数据清洗维基数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。