arXiv:2602.01276cs.AI2026-02被引 3

用大模型自动构建企业知识图谱本体,提升效率。

LLM-Driven Ontology Construction for Enterprise Knowledge Graphs

  • 分两阶段:先提取核心类与属性,再逻辑建模成层级结构。
  • 在数据领域达0.724的模糊匹配F1分数,但层次推理仍有不足。
  • 适合需要快速构建行业本体的企业和研究者。

企业知识图谱已成为统一异构数据和实施语义治理的关键工具。然而,其底层本体的构建仍依赖大量人力和领域专家知识,过程繁琐。本文提出OntoEKG,一种基于大语言模型的流水线方法,旨在从非结构化企业数据中加速生成特定领域的本体。该方法将建模任务分解为两个阶段:提取模块识别核心类与属性,推理模块则对这些元素进行逻辑结构化,形成层级关系后序列化为标准RDF格式。针对现有端到端本体构建缺乏全面评估基准的问题,我们采用来自数据、金融和物流领域的文档构建新评测数据集。实验结果表明,该方法在数据领域取得0.724的模糊匹配F1分数,展现出巨大潜力,同时也暴露出范畴定义不清和层级推理能力有限等挑战。

原文摘要 · Abstract (English)

Enterprise Knowledge Graphs have become essential for unifying heterogeneous data and enforcing semantic governance. However, the construction of their underlying ontologies remains a resource-intensive, manual process that relies heavily on domain expertise. This paper introduces OntoEKG, a LLM-driven pipeline designed to accelerate the generation of domain-specific ontologies from unstructured enterprise data. Our approach decomposes the modelling task into two distinct phases: an extraction module that identifies core classes and properties, and an entailment module that logically structures these elements into a hierarchy before serialising them into standard RDF. Addressing the significant lack of comprehensive benchmarks for end-to-end ontology construction, we adopt a new evaluation dataset derived from documents across the Data, Finance, and Logistics sectors. Experimental results highlight both the potential and the challenges of this approach, achieving a fuzzy-match F1-score of 0.724 in the Data domain while revealing limitations in scope definition and hierarchical reasoning.

知识图谱大模型本体构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。