arXiv:2512.00590cs.CLcs.AI2025-12Conference of the …被引 8

用大模型构建对齐维基数据的高质量知识图谱,提升推理准确性。

Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models

  • 分阶段构建图谱,融合维基数据类型与关系约束
  • 在MuSiQue上96%答案实体被覆盖,性能媲美带上下文的基线
  • 高效且可扩展,构建仅需不足1000个输出词元

知识图谱为大语言模型提供结构化、可验证的语义支撑,但现有基于大模型的方法多将其作为文本检索的辅助结构,未充分挖掘其内在质量。本文提出Wikontic,一种多阶段流水线,从开放域文本中提取含限定词的候选三元组,强制执行维基数据的类型与关系约束,并通过实体归一化减少重复。生成的图谱紧凑、符合本体、连接紧密;在MuSiQue数据集上,正确答案实体出现在96%的生成三元组中。在HotpotQA上,仅使用三元组的设置达到76.0 F1,MuSiQue上达59.8 F1,匹配或超越多个依赖文本上下文的检索增强生成基线。此外,在MINE-1基准上信息保留率达86%,优于以往方法。构建过程高效:图谱生成少于1000个输出词元,约为AriGraph的3倍,GraphRAG的1/20。该流程显著提升生成图谱质量,为大模型利用结构化知识提供可扩展方案。

原文摘要 · Abstract (English)

Knowledge graphs (KGs) provide structured, verifiable grounding for large language models (LLMs), but current LLM-based systems commonly use KGs as auxiliary structures for text retrieval, leaving their intrinsic quality underexplored. In this work, we propose Wikontic, a multi-stage pipeline that constructs KGs from open-domain text by extracting candidate triplets with qualifiers, enforcing Wikidata-based type and relation constraints, and normalizing entities to reduce duplication. The resulting KGs are compact, ontology-consistent, and well-connected; on MuSiQue, the correct answer entity appears in 96% of generated triplets. On HotpotQA, our triplets-only setup achieves 76.0 F1, and on MuSiQue 59.8 F1, matching or surpassing several retrieval-augmented generation baselines that still require textual context. In addition, Wikontic attains state-of-the-art information-retention performance on the MINE-1 benchmark (86%), outperforming prior KG construction methods. Wikontic is also efficient at build time: KG construction uses less than 1,000 output tokens, about 3$\times$ fewer than AriGraph and $<$1/20 of GraphRAG. The proposed pipeline enhances the quality of the generated KG and offers a scalable solution for leveraging structured knowledge in LLMs.

知识图谱大模型结构化知识维基数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。