用图结构引导大模型,自动构建可审计的领域本体。
GrOIL: Graph-Grounded Domain Ontology Induction with Constrained LLM Mediation

- 通过统一论述超图捕捉文档实体与逻辑关系,分七步生成本体。
- 在寿险领域上,关键术语覆盖率高达0.85,超过纯大模型方法。
- 每条知识都有溯源链,适合需要可解释性的领域应用。
从领域文档构建形式化本体需同时满足语料扎根、词汇一致、公理表达力和全程可追溯性,现有系统无法兼顾。本文提出七阶段图基流水线,将文档转为完整的、可审计的Web本体语言(OWL)TBox,无任何不受控生成步骤。文档首先编码为统一论述超图(UDH),捕获实体参与和论述依赖;后续阶段将图证据转化为类层次、类型化对象与数据属性、限制公理,大模型仅用于狭窄的图基中介任务。配对的断言盒(ABox)填充程序将命名个体嵌入诱导出的TBox,支持基于SPARQL的功能评估。每个生成项均带有从原始文本到各阶段的完整决策链,使本体可直接审计,适合针对性人工优化。在寿险领域使用两个基准和一个包含100份合同、覆盖十种产品类型的全新语料库进行评估,该流水线在所有评价维度表现优异:在一份定期寿险合同上,能力问题覆盖率达0.85,显著优于直接大模型(0.63)和多智能体基线(0.62);另一份合同上达0.77,高于基线(0.40和0.44);关键短语覆盖率接近人工参考本体,结构化缺口与重叠推理性能良好,且无需任何人工本体工程。本体增长分析显示词汇饱和迹象,表明该流水线能从大规模文档中产出稳定、可复用的领域表示。
原文摘要 · Abstract (English)
Constructing formal ontologies from domain documents requires simultaneously enforcing corpus grounding, vocabulary consistency, axiom-level expressivity, and end-to-end provenance, a combination no existing automatic system delivers. We present a seven-stage graph-grounded pipeline that converts domain documents into a complete, auditable Web Ontology Language (OWL) Terminological Box (TBox) without any unconstrained generation step. Documents are first encoded as Unified Discourse-Hypergraphs (UDH) capturing entity participation and discourse dependencies; subsequent stages transform this graph evidence into a class hierarchy, typed object and datatype properties, and restriction axioms, with Large Language Model (LLM) usage restricted to narrow, graph-grounded mediation tasks. A paired Assertional Box (ABox) population procedure grounds named individuals in the induced TBox, enabling SPARQL-based functional evaluation. Every emitted term carries a full decision chain from raw source passages through each pipeline stage, making the TBox directly auditable and suitable for targeted human refinement. Evaluated on the life insurance domain using two established benchmarks and a new 100-contract corpus spanning ten product types, our pipeline achieves strong results across all evaluation dimensions, outperforming direct and multi-agent LLM baselines on competency-question (CQ) coverage (0.85 vs. 0.63 and 0.62 on one term-life contract; 0.77 vs. 0.40 and 0.44 on another contract), while also attaining high keyphrase coverage comparable to a manually-constructed reference ontology and strong performance on structured gap-and-overlap reasoning, all without any manual TBox engineering. Ontology growth analysis provides evidence consistent with vocabulary saturation at scale, demonstrating that the pipeline produces stable, reusable domain representations from large document corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。