arXiv:2505.23628cs.CLcs.AI2025-05ACL被引 33

用大模型自动构建无预设模式的知识图谱,提升大模型事实准确性。

AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora

  • 利用大模型从海量文本中同时抽取三元组并动态生成完整模式
  • 构建含9亿节点、59亿边的ATLAS知识图谱,在多跳问答上超越基线
  • 无需人工干预,模式对齐人类设计达92%,适合增强大模型知识

我们提出AutoSchemaKG,一种完全自主的知识图谱构建框架,无需预先定义模式。该系统利用大语言模型从文本中同步抽取知识三元组并直接诱导全面的模式,既建模实体也建模事件,并通过概念化将实例组织为语义类别。处理超过5000万份文档后,我们构建了ATLAS(Automated Triple Linking And Schema induction)知识图谱家族,包含9亿以上节点和59亿条边。该方法在多跳问答任务中优于现有最优基线,并提升大模型的事实性。值得注意的是,我们的模式诱导与人工设计模式的语义对齐率达92%,且无需任何人工干预,证明了可动态生成模式的百亿规模知识图谱能有效补充大语言模型的参数化知识。

原文摘要 · Abstract (English)

We present AutoSchemaKG, a framework for fully autonomous knowledge graph construction that eliminates the need for predefined schemas. Our system leverages large language models to simultaneously extract knowledge triples and induce comprehensive schemas directly from text, modeling both entities and events while employing conceptualization to organize instances into semantic categories. Processing over 50 million documents, we construct ATLAS (Automated Triple Linking And Schema induction), a family of knowledge graphs with 900+ million nodes and 5.9 billion edges. This approach outperforms state-of-the-art baselines on multi-hop QA tasks and enhances LLM factuality. Notably, our schema induction achieves 92\% semantic alignment with human-crafted schemas with zero manual intervention, demonstrating that billion-scale knowledge graphs with dynamically induced schemas can effectively complement parametric knowledge in large language models.

知识图谱大模型自动构建动态模式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。