用问答对中间表示提升知识图谱构建的覆盖率与连通性
SocraticKG: Knowledge Graph Construction via QA-Driven Fact Extraction
- 通过5W1H引导的问答对扩展,先解析文档语义再提取三元组
- 在MINE和HotpotQA上实现更高事实保留率与结构连贯性
- 适合需要复杂多跳推理的知识图谱构建场景
从非结构化文本构建知识图谱(KG)可提供结构化知识表示与推理框架,但现有大模型方法面临根本性权衡:事实覆盖率高时关系碎片化,过早整合则导致信息丢失。为此,我们提出SocraticKG,一种自动化知识图谱构建方法,引入问答对作为结构化中间表示,在三元组提取前系统展开文档级语义。通过5W1H引导的问答扩展,SocraticKG捕获上下文依赖与隐含关系链,提供对源文档的显式语境支撑,有效缓解隐式推理错误。在MINE基准与HotpotQA下游任务上的评估表明,该方法能有效解决覆盖率与连通性的权衡问题,实现更优的事实保留与结构一致性,并支持复杂多跳推理。
原文摘要 · Abstract (English)
Constructing Knowledge Graphs (KGs) from unstructured text provides a structured framework for knowledge representation and reasoning, yet current LLM-based approaches struggle with a fundamental trade-off: factual coverage often leads to relational fragmentation, while premature consolidation causes information loss. To address this, we propose SocraticKG, an automated KG construction method that introduces question-answer pairs as a structured intermediate representation to systematically unfold document-level semantics prior to triple extraction. By employing 5W1H-guided QA expansion, SocraticKG captures contextual dependencies and implicit relational links typically lost in direct KG extraction pipelines, providing explicit grounding in the source document that helps mitigate implicit reasoning errors. Evaluation on the MINE benchmark and HotpotQA downstream task demonstrates that our approach effectively addresses the coverage-connectivity trade-off, achieving superior factual retention and structural cohesion while supporting complex multi-hop reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。