评测大模型零样本构建知识图谱能力,发现其能准确提取但行为差异大。
Are Large Language Models Effective Knowledge Graph Constructors?
- 分三阶段构建:初提、拆分、抽象,提升结构质量
- 在儿科心理研究数据上,三元组准确率高,幻觉少
- 公开数据集与图谱,适合医疗知识建模研究
知识图谱广泛应用于知识密集型任务,但当前大型语言模型(LLMs)在无需预设模式、零样本条件下,能否有效构建基于文档的知识图谱仍不明确。本文提出细节到抽象的层次化知识图谱(D2A-HKG)构建框架,将知识图谱构建分为初始抽取、拆分和抽象三个阶段,并从语义与结构两个角度评估生成结果。使用七种前沿大模型,在来自儿童心理健康研究论文的CMW-Lit数据集上进行零样本评测。该数据集因证据异质、因素关联复杂、关系带有统计标注而极具挑战性。结果显示,先进大模型能生成相关且忠实于原文的三元组,幻觉较少,但在不同构建阶段表现出显著不同的抽取行为。研究为前沿大模型直接构建知识图谱的能力提供了实证洞察。我们进一步公开了CMW-Lit数据集及生成的知识图谱,为后续专家精修和下游知识密集型应用提供基础资源。
原文摘要 · Abstract (English)
Knowledge graphs (KGs) are widely used in knowledge-intensive applications, yet it remains unclear how effectively current large language models (LLMs) can construct document-grounded KGs in a zero-shot, schema-free setting without relying on complex task-specific frameworks. We introduce Detail-to-Abstract Hierarchical Knowledge Graph (D2A-HKG) construction framework, which decomposes KG construction into three stages: initial extraction, splitting, and abstraction, and evaluates the resulting graphs from both semantic and structural perspectives. Using seven frontier LLMs, we benchmark zero-shot KG construction on CMW-Lit, a dataset derived from published paediatric research articles on children's mental well-being. CMW-Lit provides a challenging test bed due to its heterogeneous evidence, interconnected factors, and complex, statistically qualified relationships. Our results show that state-of-the-art LLMs can generally produce relevant and document-faithful triples with limited hallucination, while exhibiting substantially different extraction behaviors across the construction stages. These findings provide empirical insight into the strengths and limitations of frontier LLMs for direct knowledge graph construction. We further release CMW-Lit and the resulting knowledge graphs as resources for future research, with the generated graphs providing a strong foundation for expert refinement and downstream knowledge-intensive applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。