arXiv:2512.08398cs.IRcs.CL2025-12被引 1

用知识图谱解析工业标准文档中的复杂规则,提升智能问答效果。

Ontology-Based Knowledge Graph Framework for Industrial Standard Documents via Hierarchical and Propositional Structuring

  • 将文档拆解为条件与数值规则的原子命题,构建分层语义结构。
  • 在多跳问答、表格查询等任务上优于现有KG-RAG方法,准确率显著提升。
  • 适合需要精准理解工业规范的工程师与智能文档系统开发者。

基于本体的知识图谱(KG)构建是实现领域知识多维理解与高级推理的核心技术。工业标准包含大量技术信息与复杂规则,以表格、适用范围、约束、例外和数值计算等多种高度结构化形式呈现,使知识图谱构建尤为困难。本文提出一种方法,将此类文档组织为分层语义结构,将句子与表格分解为源自条件与数值规则的原子命题,并通过大语言模型(LLM)进行三元组提取,整合进本体-知识图谱中。该方法有效捕捉了文档的层级与逻辑结构,充分表达了传统方法无法反映的领域语义。为验证有效性,我们从工业标准中构建了规则、表格与多跳问答数据集,以及有毒条款检测数据集,并实现了本体感知的KG-RAG框架进行对比评估。实验结果表明,相较于现有KG-RAG方法,本方法在所有问答类型上均取得显著性能提升。研究证明,即使面对交织的条件、约束与适用范围,仍可实现可靠且可扩展的知识表示,为未来特定领域RAG发展与智能文档管理提供支持。

原文摘要 · Abstract (English)

Ontology-based knowledge graph (KG) construction is a core technology that enables multidimensional understanding and advanced reasoning over domain knowledge. Industrial standards, in particular, contain extensive technical information and complex rules presented in highly structured formats that combine tables, scopes of application, constraints, exceptions, and numerical calculations, making KG construction especially challenging. In this study, we propose a method that organizes such documents into a hierarchical semantic structure, decomposes sentences and tables into atomic propositions derived from conditional and numerical rules, and integrates them into an ontology-knowledge graph through LLM-based triple extraction. Our approach captures both the hierarchical and logical structures of documents, effectively representing domain-specific semantics that conventional methods fail to reflect. To verify its effectiveness, we constructed rule, table, and multi-hop QA datasets, as well as a toxic clause detection dataset, from industrial standards, and implemented an ontology-aware KG-RAG framework for comparative evaluation. Experimental results show that our method achieves significant performance improvements across all QA types compared to existing KG-RAG approaches. This study demonstrates that reliable and scalable knowledge representation is feasible even for industrial documents with intertwined conditions, constraints, and scopes, contributing to future domain-specific RAG development and intelligent document management.

知识图谱工业标准RAG大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。